Prompt Engineering Fundamentals: Write Instructions Models Follow
by Laura Mbeki
Build a small eval suite that tells you honestly whether a prompt change was an improvement.
LM Created by Laura Mbeki
Every bullet below is something you will have built, shipped or be able to explain by the time you finish the last lesson.
6 modules · 45 lessons · 6h of material
7 lessons running 52m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.
Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.
8 lessons running 1h 2m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.
Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.
8 lessons running 1h 4m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.
Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.
8 lessons running 1h 6m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.
Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.
7 lessons running 58m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.
Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.
7 lessons running 58m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.
Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.
6 modules · 45 lessons
6h total length
Short list, and deliberately so. If you meet these you can start today.
Every team building with models reaches the same wall. Someone edits a prompt, tries it three times, declares it better, and ships. Two weeks later nobody can say whether the system improved or quietly regressed. Evals are the cure, and they are far more approachable than the literature suggests.
You build an eval suite from nothing. First a dataset, drawn from real inputs rather than invented ones, deliberately weighted toward the cases that hurt. Then graders in increasing order of difficulty: exact match and structural checks for anything typed, similarity and rubric scoring for open text, and model-graded evaluation with all its pitfalls, including position bias, verbosity bias and graders that agree with themselves more than with reality.
You will learn to calibrate a model grader against human labels, to compute agreement honestly, and to know how many samples you need before a difference means anything. Sampling variance gets its own module, because comparing two prompts on ten examples proves almost nothing and most teams do exactly that.
The final modules put evals into a workflow: running them in continuous integration, tracking cost and latency alongside quality, catching regressions when a provider updates a model, and keeping a golden set that does not slowly leak into your prompts.
Still unsure about something? Write to misteryjj100@gmail.com and a human answers, usually the same working day.
Reviews are written by people who bought this course. We publish the critical ones too.
4.0
Rated 4.0 out of 5Course rating · 2 reviews
Alina Dobrescu
ML engineer
Being shown that a five-example comparison can flip its verdict on a rerun is the wake-up call most teams need, mine very much included. Not five, though: the continuous integration chapter is written against one particular runner, and adapting it to ours took longer than watching the chapter did.
Jasper Nieuwenhuis
Data scientist
It does not pretend a model judge is neutral. Position bias and length bias each get their own treatment, backed by measurements rather than assertions, and that candour is rare enough to justify the money by itself. The advice on writing a rubric is thinner than I hoped, and that happens to be exactly where I am stuck.
Prompt engineer and AI workflow designer
Laura went independent after seven years of agency work and now designs the prompt libraries that sit behind other people's products. She treats prompting as engineering: versioned prompts, a held-out evaluation set, a regression run before anything ships, and a token budget you have to hit. Her packs are the ones she uses with her own clients — briefing, rewriting, summarising, review — rather than sanitised examples, and each comes with notes on where it fails. She keeps every pack working across ChatGPT, Claude and a small open model, so the technique outlives the model.
One-time payment · lifetime access