Lead AI / RAG Engineer: Self-Hosted Knowledge Platform at Eventum AI
Lead AI / RAG Engineer: Self-Hosted Knowledge Platform
Eventum AI
65 to 90 USD Gross per hour
Remote
Eventum AI is an AI engineering consultancy building production-grade LLM, RAG, agent, and document-intelligence systems.
Responsibilities:
• Lead architecture and hands-on implementation of a production RAG / knowledge platform
• Build document ingestion for Google Workspace, PDFs, scans, tables, and images
• Implement permission-aware retrieval with ACL enforcement at query time
• Build hybrid search, embeddings, reranking, and citation-backed generation
• Deploy and benchmark open-weight models using vLLM or similar infrastructure
• Build evaluation/regression systems for retrieval, answer quality, security, and model upgrades
• Own production deployment, observability, CI/CD, documentation, and handoff
• Work directly with Eventum’s senior team and client technical leadership
Requirements:
• 6+ years of professional software/ML engineering experience
• Strong experience building production LLM/RAG systems
• Excellent Python/backend engineering skills
• Hands-on experience with embeddings, vector search, reranking, and document pipelines
• Experience serving open-weight LLMs with vLLM, SGLang, Triton, TGI, or similar
• Strong cloud/infrastructure skills; Azure, Docker, Terraform/Kubernetes, and CI/CD are especially relevant
• Experience with LLM evaluation and production reliability
• Comfortable owning ML, backend, data, and infrastructure rather than working in a narrow specialty
Optional:
• Google Workspace APIs and permissions
• pgvector/Qdrant
• BM25 + dense retrieval
• OCR/document AI
• multimodal models
• Prometheus/Grafana
• no-egress or security-sensitive environments
Comments
Member discussion for this news item or vacancy.