- Published on
- · Peter Yang
How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel
This post is auto-generated from a public YouTube video for personal study notes. It is not an endorsement. The transcript is machine-derived and may contain errors; refer to the original video for accuracy.
About this video
Shreya and Hamel teach the industry standard course for AI evaluations taken by 4,500+ students from OpenAI, Google, and more.
In our interview, I asked them to give live feedback on the AI evals that I built for my skills and demo how anyone can run evals with ChatGPT/Claude and their free evals skill.
Don’t miss this up to date primer on evals from the real experts. 🙂
We talked about: (00:00) Why AI evals are completely different now with the latest models (02:30) Demo: Live audit of the evals that I built for my AI skills (05:14) Top down vs. bottom up evals and why AI sucks at the latter (08:16) Spinning up AI agents to grade each eval criteria (14:32) Demo: Using Shreyas AI skill to run evals in Claude (23:46) How to review output labels to identify failures (34:37) How to turn failures into reusable evals (43:19) Where human judgement is still needed
Resources mentioned:
Thanks to our sponsors:
Where to find Shreya and Hamel: