DOIT Software
Senior AI/ML Engineer (LLMs, AI evaluation)
Summary
DOIT Software is seeking a Senior Applied AI Engineer to improve production AI systems through evaluation, experimentation, and system design. This hands-on, product-focused role combines strong ML fundamentals with the discipline of measuring and iterating on production AI systems.
What you'll do
AI Evaluation and System Quality (~30%):
- Design evaluation strategies for LLM and agent workflows
- Create metrics and KPIs for AI system performance
- Build and maintain evaluation datasets
- Debug production AI failures systematically
Prompt and Agent System Design (~30%):
- Improve agent orchestration and workflows
- Diagnose failures across agent pipelines
- Refine system prompts and agent interactions
- Improve reliability, latency, and response quality
ML and AI Systems (~30%):
- Contribute to recommendation systems, ranking, and personalization
- Work on itinerary optimization and constraint-based planning
- Improve LLM-based reasoning systems
- Optional computer vision pipelines
Engineering Integration (~10%):
- Collaborate with backend engineers using Golang, Python, Postgres, and Redis
Requirements
Must-haves:
- Strong AI/ML fundamentals including evaluation metrics (precision/recall/F1), ranking and recommendation concepts, embeddings and similarity, experimentation methodology
- Evaluation-driven mindset with ability to think in metrics, design experiments, measure improvements quantitatively, and debug failures methodically
- Experience with LLM systems including prompt design, agent workflows, evaluation of LLM outputs, and production LLM integrations
- Ability to ship production systems and iterate based on results
- Programming ability in Python, Go, or similar languages with comfort reading and modifying production code
Strong Signals (Nice to Have):
- Experience improving an AI system after deployment
- Recommendation systems or ranking experience
- Optimization or constraint-based systems experience
- Computer vision experience
- Experience building evaluation frameworks
- Golang experience
- Startup or small-team engineering experience
Conditions
This is a hands-on, product-focused role with collaborative work alongside the AI platform lead. The role involves working on systems that real users depend on, requiring the ability to balance exploration with delivery and work with partially-defined problems. Not a fit for those seeking research-focused roles without production deployment, those who rely heavily on frameworks without understanding fundamentals, or those preferring narrow specialization over product ownership.