Loading article
Putting language models into products people rely on: keeping generated content grounded, designing rubrics that produce evidence, and knowing when the model is the wrong tool.
I like LangGraph and I have removed it from a project. The four things a graph actually buys you, and the test I use before reaching for one.
RAG is easy to demo and hard to trust. What it took to make an AI curriculum generator produce lessons I would put in front of a student preparing for CSIR NET.
Building LLM-assisted interview assessment: rubrics that produce evidence instead of verdicts, the variance nobody reports, and why transcription quality is a fairness problem.
All articles in chronological order, or get in touch if this is the kind of problem you’re working on.