← ProjectsMini Calc TA
  • Python
  • Anthropic Claude API
  • Pydantic
  • YAML
← All Projects

2026 · Solo

Mini Calc TA

An agentic LLM grading pipeline that checks a student's calculus work step by step, plus a six-configuration ablation study that made me reverse my own shipped architecture.

  • Python
  • Anthropic Claude API
  • Pydantic
  • YAML

The idea

Grading a calculus solution isn’t a single yes/no. It’s a chain of steps, and the whole point is finding where it broke, not just whether it broke. So Mini Calc TA verifies each step of a student’s solution in isolation, locates the first incorrect step, and assigns partial credit from there.

Isolating each step matters more than it sounds like it should. Show the model the final answer alongside the work, and it’ll happily anchor on “is the answer right?” and retroactively justify wrong reasoning to get there. Per-step isolation keeps it honest.

Building the harness before the agent

Before writing a single line of agent code, I froze a 12-case eval harness and a deterministic scorer. That order matters. It meant every architecture I tried afterward got graded against the exact same yardstick, not vibes.

Then I ran a six-configuration ablation, 4–5 independent trials each. The winning architecture cut grading error 40% against a single-prompt baseline (MAE 0.87 → 0.52) at 97% first-error-location accuracy. Genuinely happy with that number!

Where it gets humbling

Here’s the part I’m most proud of, and it’s not the 40%. On repeat trials, I noticed the design I’d actually shipped had an error range that fully overlapped the baseline’s (Mann-Whitney U, p ≈ 0.008 for the replacement). Statistically, it wasn’t beating the thing it was supposed to replace.

Digging in, I found a silent token-truncation failure quietly producing plausible-looking garbage scores, with no error raised anywhere. No crash, no red flag, just confidently wrong numbers. So I reversed the shipped architecture in favor of the one the ablation actually supported. Not the fun answer, but the correct one.

What it demonstrates

That an eval harness isn’t a formality you bolt on after the fact. It’s the thing that catches you when your own architecture is quietly lying to you. Building the test before the system it tests is what turned a plausible-looking bug into a caught one.