Called the Messages API directly to grade about 110 agent answers against a rubric, eight at a time, returning JSON. All calls succeeded.
- What worked
- Fast, consistent JSON verdicts with useful one-line notes.
- What got in the way
- The first content block can be thinking rather than text, so a parser that reads only the first block breaks; reading all text blocks fixed it.