Senior AI Engineer (LLMs & Agents)
Software Engineering, Data Science · Full-time
Austin, TX, USA · Remote
About Cooklist
Cooklist is the AI grocery-intelligence platform powering meal planning and shopping for millions of consumers across our consumer app and white-label enterprise suite. Our mission is to combine the intelligence of a personal shopper, chef, and nutritionist to help people save time, eat better, and enjoy happier lives.
We're profitable, process billions of dollars in transactions at the nation's largest retailers, and our mobile experiences reach millions of people. We're backed by Techstars, Mercury Fund, and industry leaders including the former Chief Technology Officer of a leading U.S. grocer and the former Chief Product Officer of Amazon Fresh.
Role Overview
We’re hiring a Senior AI Engineer (LLMs & Agents) to own the intelligence behind the AI shopping assistants Cooklist has built for leading US grocers and the Cooklist app. You will improve the prompts, tools, retrieval, policies, evaluations, and workflow architecture that make an agentic system trustworthy at scale.
Success is measured through agent reliability, shopper engagement, and retention. You will connect model and workflow improvements to those outcomes through rigorous evaluation, monitoring, and product experimentation. This is a high‑ownership role on a small team where your decisions directly impact millions of shoppers, where accuracy around allergens and nutrition is critical.
We are an AI-leveraged engineering organization. Our question is always "How do we use AI to build AI?" You will establish patterns, tests, and guardrails that allow AI assistants to safely contribute to the codebase and compound our output.
Responsibilities
Own LLM reliability end‑to‑end: architect prompts, tools, and reasoning workflows that meet strict accuracy, safety, and latency requirements.
Design robust evals: build offline/online eval suites for structured output, factuality, grounding, allergen sensitivity, and user‑goal attainment; define gold sets, synthetic data pipelines, and automatic failure taxonomies.
Productionize agent workflows: retrieval‑augmented generation, tool calling over GraphQL/WebSockets, function/tool schemas, and strict JSON output contracts.
Model strategy: evaluate and deploy model mixes (reasoning vs. fast paths), caching strategies, and guardrails to balance quality, latency, and cost.
Monitoring & observability: ship real‑time conversation analytics, drift detection, canary/shadow testing, incident taxonomies, and auto‑triage for misbehavior.
Protect Shopper Trust: encode safeguards for allergens, dietary restrictions, nutrition, product availability, price, promotions, and other grocery-specific constraints.
Tight product loop: partner with mobile/backend to ship agent features, collect outcome‑level telemetry, and iterate quickly (“build first, refine fast”).
Scale the system: turn customer learnings into reusable libraries, datasets, tests, and playbooks that make each future retailer deployment faster and safer.
Qualifications
You’ve shipped LLM systems to production with real user impact. Ideally agentic loops, tool calling, and structured outputs at scale.
You’re fluent in Python and have built eval harnesses, automated datasets, and dashboards for LLM quality.
You’ve implemented RAG (indexing, chunking, embeddings, reranking) and understand failure modes (hallucination, grounding, duplication, drift).
You can design and enforce strict schemas, guarantee parseability, and create deterministic fallbacks.
You’re comfortable making model tradeoffs (reasoning models vs. smaller/cheaper paths; latency budgets; cost controls) and can prove the impact.
You care about safety (allergens, dietary needs, policy adherence) and can translate product risk into tests, gates, and roll‑out controls.
You move with founder energy: high ownership, high bar for polish, gritty, and calm under production pressure.
Our Stack
- Language: Python/Django backend; Javascript/React Native frontend
- APIs/Data: GraphQL; real‑time streaming over WebSockets
- Mobile: React Native (close collaboration with the mobile team)
- LLM engineering: internally built prompt/tool libraries, RAG pipelines & eval system
What We Offer
- Competitive compensation + meaningful equity
- Austin, TX based with WFH flexibility
- Work directly with founders and an elite, tight‑knit team
- Ship experiences that materially improve the lives of millions
- A high‑intensity, high‑ownership environment designed for builders