hireclover
All jobs

PandaDoc

Machine Learning Engineer - Document Intelligence & Applied GenAI

Remote Senior $4.7k–$12.8k/moest.

Summary

PandaDoc is hiring a Machine Learning Engineer to build and optimize models for document intelligence and applied GenAI. You will develop evaluation frameworks, create datasets, train transformer-based models, and deploy ML systems that enable intelligent document understanding and extraction at scale.

What you'll do

  • Model Development & Evaluation: Build evaluation frameworks for document models, LLMs, OCR, and structured extraction; define metrics, benchmarks, and validation strategies.
  • Dataset & Pipeline Creation: Design and curate high-quality datasets for training and fine-tuning; create scalable preprocessing pipelines for PDFs, scans, images, forms, and semi-structured documents.
  • Model Training & Fine-Tuning: Train and fine-tune transformer-based OCR, VLMs, layout models, and open-source LLMs for document understanding; optimize for reliability, accuracy, and cost efficiency.
  • Inference & Deployment: Deploy ML models using modern inference runtimes (vLLM, TGI, TensorRT, ONNX Runtime); build guardrails, monitoring, and fallback mechanisms.
  • RAG & Document Reasoning: Develop retrieval and chunking strategies for document structures; optimize end-to-end RAG pipelines for semantic search and workflow automation.
  • Cross-Functional Collaboration: Partner with PMs, backend engineers, and product designers to define AI opportunities and translate requirements into technical solutions.

Requirements

  • Required: 5+ years of Python experience
  • Experience training, fine-tuning, and deploying computer vision models for document intelligence tasks (layout detection, table extraction, OCR, information extraction)
  • Hands-on experience with document understanding frameworks: LayoutLM, Donut, DocFormer, DeepSeek-OCR, LightOnOCR-1B
  • Experience deploying and optimizing models using vLLM, TGI, TensorRT, or ONNX Runtime
  • Experience applying LLMs to document intelligence workflows, including frontier and open-source models
  • Strong understanding of coordinate systems and spatial reasoning for field detection in forms and documents
  • Nice to Have: Familiarity with PDF parsing libraries and document preprocessing pipelines; experience fine-tuning open-source models for domain-specific tasks; knowledge of evaluation metrics for document understanding (F1, exact match, etc.)

Conditions

  • Remote work opportunity — distributed team worldwide (Lisbon to Manila, Florida to California)
  • 6 self-care days
  • Competitive salary
  • Open culture emphasizing feedback and professional development
  • Additional benefits available

Browse jobs