ai
3 мин
26 сентября 2026 г.
Источник: Dev.to AI Feed

My AI feature was failing 26% of the time. Nothing looked broken.

Angelina Gupta
Angelina Gupta
RSS AI Ingest
My AI feature was failing 26% of the time. Nothing looked broken.

I built TaskFlow, a project management app with an AI feature called Quick Add. You type "assign the API docs to Priya by Friday" and it creates a task with a title, an assignee, and a due date. For twelve days in August, it quietly stopped...

I built TaskFlow, a project management app with an AI feature called Quick Add. You type "assign the API docs to Priya by Friday" and it creates a task with a title, an assignee, and a due date. For twelve days in August, it quietly stopped doing that for about a quarter of requests. Users got a task titled "assign the API docs to Priya by Friday" — no assignee, no date. No error. No crash. Just a worse product. My test suite caught it on day one. I didn't look until day twelve. This post is about both of those things. Why LLM features need their own tests A normal unit test checks that a function returned. It can't tell you whether the model got the date right. LLM output looks fluent even when it's wrong, so "it didn't crash" means very little. So I built an eval harness: 71 test cases across 4 prompt suites — quick-add (50), extract-tasks (9), decompose (6), and today (6). Each case is an input plus the fields the model should return. A few decisions that turned out to matter: It imports the real production prompts, not copies. If a prompt changes, the tests test the change. Scoring is deterministic. No LLM grading another LLM — dates and assignees either match or they don't. The clock is pinned, so "Friday" means the same date on every run. It mirrors production's token limits. This one ends up being the whole story. It runs in GitHub Actions: a small subset on every pull request, the full suite nightly. It separates a wrong answer from a failed request. Rate-limit errors get retried, so a 429 doesn't show up as a quality failure. What happened Groq retired the model I was using, llama-3.3-70b-versatile, on August 16. I switched to openai/gpt-oss-120b on August 24. The nightly suite went red that same night. It stayed red for twelve consecutive nights before I looked properly. Two things kept it hidden: The fallback worked too well. When Quick Add couldn't parse the model's output, it used the raw text as the task title. That's the right design — the app stays usable — but it meant there was no error to notice. The error looked like someone else's problem. 13 of the 15 failures were 400 Failed to validate JSON. That reads like an infrastructure issue, so my first instinct was to blame the provider. Finding the cause The error message was a dead end. The distribution wasn't. Suite Result max_tokens quick-add 35/50 300 extract-tasks 9/9 1500 decompose 6/6 1800 today 6/6 900 Every failure was in the one suite with the smallest token limit. The token data confirmed it. Of the quick-add cases that returned, the five largest completions were 283, 284, 288, 298 and 298 tokens — jammed right up against the 300 limit. The old model's largest was 64. The new model is a reasoning model. It spends tokens thinking before it writes any JSON. With a 300-token cap, it was running out mid-object and returning truncated JSON, which the provider then rejected. The fix — and what it uncovered The first fix was one number: max_tokens from 300 to 900, in two places that must always match — the production controller and the eval harness. That exposed three more bugs the 400s had been hiding. Once the model could finish its answers, three date cases returned due: null: "by end of week" "by July 20" "in two weeks" The prompt gave the model a 10-day calendar and told it never to compute weekdays itself. So anything outside that window came back empty. The proof: "in 3 days" and "by July 9" already passed, because both fall inside it. I split the rule. Weekday phrases still resolve from the calendar; absolute dates and offsets are computed forward from today. The prompt now says the calendar is "a 10-day window, not a limit on what you may answer." Priority had a similar gap — it only recognised words like "asap" or "critical", so "this is blocking the release" came back as unstated. It now judges described impact, not tone. Quick-add went from 35/50 to 50/50. The full suite went from 78.9% to 98.6%. Where it stands today As of the latest nightly, it's 69/71 — 97.2%. Two cases fail. One is a genuine quality miss. The "today" planner includes a task it should filter out. I haven't fixed it because it's a ranking judgment, not a rule I can state cleanly — and I'd rather keep one honest failure than overfit the prompt to a single test case. The other is the same bug coming back. Same truncated-JSON signature, now in the "today" suite at its 900-token cap. My fix was too narrow: I raised the limit where it was failing, instead of asking which other endpoints had limits sized for a non-reasoning model. One endpoint is still at 400 tokens with no eval coverage at all. What I'd tell other developers Graceful fallbacks hide bugs. Make your fallback path loud — log it, count it, alert on it. Read the distribution, not the error. The error text pointed at the provider. The pattern across suites pointed at the real cause. When you fix a limit, audit every place with the same assumption. A model change breaks assumptions everywhere at once. A red test nobody reads isn't a test. The harness did its job on night one. The gap was mine. TaskFlow is live at taskflow-dpsa.vercel.app. Code: github.com/angelina10504/taskflow.

Хотите внедрить ИИ в ваш бренд?

Спроектируем и развернем автономных агентов и современный цифровой стек под ваши задачи.

Рассчитать проект