Artificial Intelligence · Remote

LLM Engineer

Full-time · 4-6 years experience

Overview

We are hiring an LLM Engineer to build the infrastructure that makes language models reliable, observable, and cost-efficient in production. You will own the model-serving layer, fine-tuning pipelines, and evaluation frameworks.

Responsibilities

  • Design and deploy model-serving infrastructure with low-latency inference and intelligent caching
  • Build fine-tuning pipelines using LoRA, QLoRA, and full-parameter tuning strategies
  • Implement evaluation frameworks that measure factual accuracy, safety, and alignment
  • Develop monitoring systems for cost tracking, token usage, and model drift
  • Architect RAG systems that ground model outputs in verified knowledge bases

Basic Qualifications

  • 4+ years in machine learning engineering with 2+ years focused on LLMs
  • Deep understanding of transformer architectures, attention mechanisms, and scaling laws
  • Proficiency with PyTorch and the Hugging Face ecosystem
  • Experience deploying models on cloud infrastructure (AWS SageMaker, GCP Vertex AI)
  • Production experience with vector databases and embedding pipelines

Preferred Qualifications

  • Published research or contributions to NLP/ML conferences
  • Experience with model quantization, distillation, and on-device deployment
  • Background in information retrieval or knowledge graph construction
  • Experience with RLHF or preference optimisation techniques

Technologies

PyTorchHugging FaceAWS SageMakerLangChainWeaviatePineconevLLM

Compensation & Benefits

  • Competitive salary commensurate with experience and location
  • Remote-first work environment with flexible scheduling
  • Annual professional development budget for conferences, courses, and tools
  • Health, dental, and vision insurance coverage
  • Equity participation in the firm's growth
  • Equipment stipend for your home office setup