PandaDoc
AI Engineer - Document Intelligence & Applied GenAI
Remote Middle $2.5k–$8.4k/moest.
Summary
PandaDoc is seeking an AI Engineer specializing in Document Intelligence and Applied GenAI. In this role, you will build and maintain evaluation frameworks for document models, develop datasets and pipelines for document understanding, and deploy ML models for real-world document processing at scale. You will work on cutting-edge GenAI workflows combining traditional document AI with vision-language models and LLM-based extraction.
What you'll do
- Model Development & Evaluation: Build evaluation frameworks for document models, LLMs, OCR, and structured extraction; define metrics, benchmarks, and validation strategies for real-world workloads
- Dataset & Pipeline Creation: Design and curate high-quality datasets for supervised training and fine-tuning; create scalable preprocessing pipelines for PDFs, scans, images, forms, and semi-structured documents
- Model Training & Fine-Tuning: Train and fine-tune transformer-based OCR, VLMs, layout models, and open-source LLMs for document understanding tasks; optimize for reliability, accuracy, and cost efficiency
- Inference & Deployment: Deploy ML models with modern inference runtimes (vLLM, TGI, TensorRT, ONNX Runtime); build guardrails and monitoring mechanisms
- RAG & Document Reasoning: Develop retrieval and chunking strategies for document structures; optimize RAG pipelines for semantic search, Q&A, and workflow automation
- Cross-Functional Collaboration: Partner with PMs, backend engineers, and product designers to define AI opportunities and translate requirements into technical solutions
Requirements
- Required: 5+ years Python experience
- Experience training, fine-tuning, and deploying traditional computer vision models for document intelligence (layout detection, table extraction, OCR, information extraction)
- Hands-on experience with document understanding frameworks: LayoutLM, Donut, DocFormer, DeepSeek-OCR, LightOnOCR-1B
- Experience deploying and optimizing models using vLLM (preferred), TGI, TensorRT, or ONNX Runtime
- Experience applying LLMs to document intelligence workflows, including frontier and open-source models
- Strong understanding of coordinate systems and spatial reasoning for absolute positioning and field detection in forms/documents
- Nice to have: Familiarity with PDF parsing libraries and document preprocessing pipelines; experience fine-tuning open-source models for domain-specific tasks; knowledge of evaluation metrics for document understanding (F1, exact match, etc.)
Conditions
- Remote-friendly: work from anywhere with a distributed team across Lisbon, Manila, Florida, California, and beyond
- 6 self-care days
- Competitive salary
- Open culture emphasizing feedback and professional/personal development