Artificial Intelligence · Remote
LLM Engineer
Full-time · 4-6 years experience
Overview
We are hiring an LLM Engineer to build the infrastructure that makes language models reliable, observable, and cost-efficient in production. You will own the model-serving layer, fine-tuning pipelines, and evaluation frameworks.
Responsibilities
- Design and deploy model-serving infrastructure with low-latency inference and intelligent caching
- Build fine-tuning pipelines using LoRA, QLoRA, and full-parameter tuning strategies
- Implement evaluation frameworks that measure factual accuracy, safety, and alignment
- Develop monitoring systems for cost tracking, token usage, and model drift
- Architect RAG systems that ground model outputs in verified knowledge bases
Basic Qualifications
- 4+ years in machine learning engineering with 2+ years focused on LLMs
- Deep understanding of transformer architectures, attention mechanisms, and scaling laws
- Proficiency with PyTorch and the Hugging Face ecosystem
- Experience deploying models on cloud infrastructure (AWS SageMaker, GCP Vertex AI)
- Production experience with vector databases and embedding pipelines
Preferred Qualifications
- Published research or contributions to NLP/ML conferences
- Experience with model quantization, distillation, and on-device deployment
- Background in information retrieval or knowledge graph construction
- Experience with RLHF or preference optimisation techniques
Technologies
PyTorchHugging FaceAWS SageMakerLangChainWeaviatePineconevLLM
Compensation & Benefits
- Competitive salary commensurate with experience and location
- Remote-first work environment with flexible scheduling
- Annual professional development budget for conferences, courses, and tools
- Health, dental, and vision insurance coverage
- Equity participation in the firm's growth
- Equipment stipend for your home office setup
