PandaDoc
Machine Learning Engineer - Document Intelligence & Applied GenAI
Remote Senior $4.7k–$12.8k/moest.
Summary
PandaDoc is seeking a Machine Learning Engineer specialized in document intelligence and applied GenAI. The role focuses on building and optimizing ML models for document understanding, including vision-language models, OCR systems, and LLM-based extraction and reasoning to directly impact customer workflows.
What you'll do
- Model Development & Evaluation: Build and maintain evaluation frameworks for document models, LLMs, OCR, and structured extraction. Define metrics, benchmarks, and validation strategies for real-world document workloads.
- Dataset & Pipeline Creation: Design and curate high-quality datasets for supervised training, fine-tuning, and validation. Create scalable preprocessing pipelines for PDFs, scans, images, forms, and semi-structured documents.
- Model Training & Fine-Tuning: Train and fine-tune transformer-based OCR, VLMs, layout models, and open-source LLMs for document understanding tasks. Optimize models for reliability, accuracy, and cost efficiency.
- Inference & Deployment: Deploy ML models with modern inference runtimes (vLLM, TGI, TensorRT, ONNX Runtime). Build guardrails, monitoring, and fallback mechanisms.
- RAG & Document Reasoning: Develop retrieval and chunking strategies for document structures. Optimize end-to-end RAG pipelines for semantic search, Q&A, and workflow automation.
- Cross-Functional Collaboration: Partner with PMs, backend engineers, and product designers to define AI opportunities and translate requirements into technical solutions.
Requirements
- 5+ years of Python experience
- Experience training, fine-tuning, and deploying computer vision models for document intelligence tasks (layout detection, table extraction, OCR, information extraction)
- Hands-on experience with document understanding frameworks: LayoutLM, Donut, DocFormer, DeepSeek-OCR, LightOnOCR-1B
- Experience deploying and optimizing models using inference frameworks: vLLM (preferred), TGI, TensorRT, or ONNX Runtime
- Experience applying LLMs to document intelligence workflows, including frontier and open-source models
- Strong understanding of coordinate systems and spatial reasoning for absolute positioning and field detection in forms/documents
- Nice to have: Familiarity with PDF parsing libraries, experience fine-tuning open-source models for domain-specific tasks, knowledge of evaluation metrics for document understanding (F1, exact match)
Conditions
- Remote work opportunity — distributed team worldwide (Lisbon to Manila, Florida to California)
- 6 self care days
- Competitive salary
- Open culture emphasizing feedback and professional development
- Additional benefits included