AI Quality

14 articles on AI Quality — programs, tooling, and delivery on Petralian.

Research papers on a desk morphing into a single clean brief stack under a desk lamp.
Hybrid

Grok 4.5 in Cursor for Knowledge Work: Beyond the Benchmark Row

Best forAnyone who read Grok 4.5 CursorBench numbers and wants a practical rule for knowledge work, briefs, and synthesis without changing every default

Grok 4.5 scores high on CursorBench agent tasks. I use it for synthesis, briefs, and research passes — not as a default for every repo edit. Here is the decision frame.

Continue reading →
Four workshop trays on a concrete bench under colored gel lights, each suggesting a different work mode, editorial still life, no logos or readable text.
Hybrid

Best Cursor Model by Work Mode (2026): Analysis, Review, Execution, Greenfield

Best forAnyone choosing Cursor model defaults by work mode who wants cost-aware picks from public benchmarks

CursorBench 3.2 reports one score per model, but agent work varies by risk and scope. Here is a work-mode default map for anyone choosing Cursor models — with cost, tokens, and steps from the public table.

Continue reading →
Green pedestrian signal beside a taller red emergency beacon on a concrete wall at dusk, shallow depth of field, no logos or readable text.
Hybrid

When to Escalate from Composer 2.5 to Fable 5: A Decision Tree

Best forAnyone governing Cursor model spend who needs escalation triggers before premium tiers become default

Composer 2.5 is the CursorBench budget default. Fable 5 tiers buy peak score at higher cost. Use this decision tree to escalate only when failure cost justifies the line item — for solo work or team policy.

Continue reading →
Five nested brass rings on dark slate with coin stacks beside each ring, macro editorial still life, amber keylight, no logos or readable text.
Hybrid

Fable 5 Pricing on Cursor: Every Tier Explained (Max to Low)

Best forAnyone approving Cursor AI spend who needs Fable tier unit economics before picking a default

Fable 5 ships as five effort tiers on Cursor. CursorBench 3.2 shows how score, cost, tokens, and steps change from Max to Low — for anyone approving model spend, not pickers chasing rank.

Continue reading →
Three seedlings sprouting from cracked concrete under distinct colored light filters, morning mist, macro editorial, no logos or readable text.
Hybrid

Open Models on CursorBench 3.2: Grok 4.5, GLM 5.2, Kimi K2.7, and LongCat

Best forAnyone comparing open-model vendor claims to Cursor session economics before changing defaults

Open-model launch posts cite SWE-bench; CursorBench cites session cost. Here is how to read Grok, GLM, Kimi, and LongCat for buying decisions — not picker hype alone.

Continue reading →
Cinematic 16:9: spreadsheet notebook beside a CI pipeline light and
Hands-on

Measure Your Cursor Harness — CSV, CI, and OpenRouter Dollars

Best forProgram leads measuring whether a Cursor harness improves output and spend

Do not build Phase 2 orchestration until Phase 0 data says so. Layer 4 feedback — CSV, footer Agents line, eval gate — plus weekly OpenRouter checks beat benchmark leaderboard anxiety.

Continue reading →
Cinematic 16:9: a single Composer pane on a workbench surrounded by
Hands-on

You Already Have an AI Harness in Cursor

Best forPractice leads governing Cursor Agent with harness discipline without microservice overhead

Terminal-Bench harnesses look like separate products. On a production Shopify app I already had subagents, CI gates, and session rules. You keep model and mode control — the harness supports routing, tests, and memory gates, not autopilot.

Continue reading →
Cinematic 16:9 macro photograph: scatter-plot points carved as glowing
Hybrid

CursorBench 3.2: Fable 5 Tops the Chart, but Composer 2.5 Wins the Budget

Best forPractice leads and commercial operators setting Cursor AI model policy using CursorBench unit economics

Fable 5 Max leads CursorBench 3.2 at 70.5%, but at 17 USD per task and 72 steps. Grok 4.5 High scores 66.7% at 1.51 USD. Composer 2.5 still wins score per dollar at 56.1% and 0.44 USD.

Continue reading →
Editorial 16:9 illustration: browser DOM tree with a hidden instruction
Hands-on

Capturing UI Designs for AI Agents Creates a Prompt Injection Surface

Best forBuilders feeding UI context to agents who want to understand prompt-injection risk

Design capture CLIs that dump outerHTML into SKILL.md files can smuggle instructions. Sanitize at the trust boundary before agents read the DOM.

Continue reading →
A transparent engineering control room with six illuminated quality
Hands-on

How We Built Gravio’s Scoring Engine: From Repo Signals to Release Gates

Best forBuilders who want the architecture behind an AI quality scoring engine

A practical breakdown of how Gravio turns repository signals into six-dimension scores, hard quality gates, and actionable remediation plans.

Continue reading →
A CI pipeline diagram where one stage is AI Quality Gate with pass/fail
Hands-on

The New CI Gate: Failing Builds on Agent Quality

Best forBuilders wiring AI quality checks into CI and release pipelines

Unit tests catch code failures. They do not always catch AI quality regressions. Here is how to add quality thresholds as a first-class release gate.

Continue reading →
A network map of many software repositories connected to one quality
Hybrid

Team Playbook: Rolling Out Gravio Across Multiple Repositories

Best forPlatform and engineering leads rolling AI quality scoring across multiple repos

A practical rollout framework for introducing Gravio across many repos without creating process fatigue, policy confusion, or noisy quality signals.

Continue reading →
A timeline dashboard with quality score trend lines bending downward
Hybrid

Why AI Agent Output Quality Drifts Over Time (And How to Catch It Early)

Best forTeams running agent workflows who need a practical quality signal before drift becomes production risk

Your AI outputs can look great this month and degrade next month without obvious failures. Here is why drift happens and how to detect it before it reaches production.

Continue reading →
A cinematic workstation scene with encrypted data streams flowing from
Hands-on

Zero-Knowledge AI Quality: How Gravio Scores Agents Without Seeing Your Code

Best forBuilders exploring privacy-preserving AI quality scoring with Gravio

Most AI quality platforms ask you to trust them with your source code. Gravio takes a different path: encrypted scoring designed to keep plaintext out of the server path.

Continue reading →