mirror of
https://github.com/stack-auth/stack.git
synced 2026-07-20 21:29:36 +08:00
## Problem
The QA reviewer (`apps/backend/src/lib/ai/qa-reviewer.ts`) extracted the
model's JSON verdict from free-form text with a greedy `/\{[\s\S]*\}/`
regex followed by `JSON.parse`. Whenever Haiku wrapped its JSON in prose
or tool-call text, the captured slice was malformed and `JSON.parse`
threw:
```
SyntaxError: Expected property name or '}' in JSON at position 1 (line 1 column 2)
```
This is Sentry issue **STACK-BACKEND-18F** — **218 occurrences in the
last 30 days** (262 all-time), still ongoing. It's caught (background QA
task, 0 users impacted) but noisy, and on failure the review is silently
marked failed.
### Root cause
The brace-matching heuristic assumes the first `{` in the response
starts the JSON object and the last `}` ends it. Any stray brace in
surrounding prose (a preamble like `"Here's my review {note}: {...}"`,
unquoted/single-quoted keys, or `<function_calls>` tool text) breaks the
extraction. The reviewer output is non-deterministic, so a low steady
fraction of reviews fail.
## Fix
Switch to schema-enforced **structured output** via `generateText`'s
`output: Output.object({ schema })` (AI SDK v6), which keeps the
existing tool-calling loop while constraining the final result to a
validated object. The verdict is read directly off `result.output`,
removing the regex, `JSON.parse`, and manual `typeof` shape validation.
- Net **−12 lines**.
- Trimmed the now-redundant "respond with ONLY valid JSON … {inline
schema}" block from the system prompt (the schema enforces structure);
kept the semantic guidance (flag types, scoring, `needsHumanReview`
rule).
- Kept the existing `overallScore` clamp.
## Verification
- **Typecheck** (`tsc --noEmit`): clean.
- **Lint** (`eslint`): clean.
- **Live end-to-end**: ran the real Haiku-4.5 reviewer against the exact
prompt/input that failed **11/15** on the old path → **15/15** return a
validated object on the new path, zero parse failures.
## Follow-ups (out of scope, noted during investigation)
- These Sentry events leak a live access token + publishable key into
the request-context payload; the OpenRouter key sits in plaintext
(commented) in `apps/backend/.env.local`. Worth rotating/scrubbing.
- Consider capturing `result.text` on the failure path so any residual
structured-output failures are inspectable (currently discarded).
<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Replaced regex/JSON.parse extraction in the QA reviewer with
schema-enforced structured output using `ai`’s `generateText` +
`Output.object`. Also defaulted `improvementSuggestions` to an empty
string to avoid false failures when the model omits suggestions.
- **Bug Fixes**
- Added a `zod` verdict schema and read result from `result.output`.
- Removed greedy JSON extraction and manual shape checks.
- Defaulted `improvementSuggestions` to "" to tolerate omission and
prevent false `needsHumanReview` verdicts.
- Slimmed prompt by dropping “respond with ONLY JSON” block; kept review
guidance.
- Kept `overallScore` clamp; tool-calling loop unchanged.
- Verification: typecheck/lint clean; 15/15 successful runs on a
previously failing case.
<sup>Written for commit
|
||
|---|---|---|
| .. | ||
| prisma | ||
| scripts | ||
| src | ||
| .env | ||
| .env.development | ||
| .eslintrc.cjs | ||
| .gitignore | ||
| instrumentation-client.ts | ||
| LICENSE | ||
| next.config.mjs | ||
| package.json | ||
| prisma.config.ts | ||
| tsconfig.json | ||
| vercel.json | ||
| vitest.config.ts | ||
| vitest.setup.ts | ||