hireclover
All jobs

PandaDoc

Machine Learning Engineer - Document Intelligence & Applied GenAI

Remote Senior $4.7k–$12.8k/moest.

Summary

PandaDoc is seeking a Machine Learning Engineer specialized in document intelligence and applied GenAI. The role focuses on building and optimizing ML models for document understanding, including vision-language models, OCR systems, and LLM-based extraction and reasoning to directly impact customer workflows.

What you'll do

  • Model Development & Evaluation: Build and maintain evaluation frameworks for document models, LLMs, OCR, and structured extraction. Define metrics, benchmarks, and validation strategies for real-world document workloads.
  • Dataset & Pipeline Creation: Design and curate high-quality datasets for supervised training, fine-tuning, and validation. Create scalable preprocessing pipelines for PDFs, scans, images, forms, and semi-structured documents.
  • Model Training & Fine-Tuning: Train and fine-tune transformer-based OCR, VLMs, layout models, and open-source LLMs for document understanding tasks. Optimize models for reliability, accuracy, and cost efficiency.
  • Inference & Deployment: Deploy ML models with modern inference runtimes (vLLM, TGI, TensorRT, ONNX Runtime). Build guardrails, monitoring, and fallback mechanisms.
  • RAG & Document Reasoning: Develop retrieval and chunking strategies for document structures. Optimize end-to-end RAG pipelines for semantic search, Q&A, and workflow automation.
  • Cross-Functional Collaboration: Partner with PMs, backend engineers, and product designers to define AI opportunities and translate requirements into technical solutions.

Requirements

  • 5+ years of Python experience
  • Experience training, fine-tuning, and deploying computer vision models for document intelligence tasks (layout detection, table extraction, OCR, information extraction)
  • Hands-on experience with document understanding frameworks: LayoutLM, Donut, DocFormer, DeepSeek-OCR, LightOnOCR-1B
  • Experience deploying and optimizing models using inference frameworks: vLLM (preferred), TGI, TensorRT, or ONNX Runtime
  • Experience applying LLMs to document intelligence workflows, including frontier and open-source models
  • Strong understanding of coordinate systems and spatial reasoning for absolute positioning and field detection in forms/documents
  • Nice to have: Familiarity with PDF parsing libraries, experience fine-tuning open-source models for domain-specific tasks, knowledge of evaluation metrics for document understanding (F1, exact match)

Conditions

  • Remote work opportunity — distributed team worldwide (Lisbon to Manila, Florida to California)
  • 6 self care days
  • Competitive salary
  • Open culture emphasizing feedback and professional development
  • Additional benefits included

Browse jobs