Sign InOpen Brain
VercelEngineering PostOfficial Source

Pixel Canary is now available in stealth for free on AI Gateway

Pixel Canary reaches 90.3% on Vercel’s Next.js eval and 96.8% with docs in AGENTS.md, but lacks ZDR and may use prompts and responses for training.

Vercel · Sep 25, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Pixel Canary is available as **stealth/pixel-canary** through Vercel AI Gateway. It passed **28 of 31 tasks (90.3%)** without supplied docs and **30 of 31 (96.8%)** when Next.js documentation was provided through AGENTS.md.

Practical Implication

For Next.js agents, test whether targeted framework documentation improves results in your own repository. The model covers migrations, data fetching, optimization, caching, and view transitions, and Vercel’s CLI can configure supported agents to use the gateway.

Agent-Ready Context
Pixel Canary is available as **stealth/pixel-canary** through Vercel AI Gateway. It passed **28 of 31 tasks (90.3%)** without supplied docs and **30 of 31 (96.8%)** when Next.js documentation was provided through AGENTS.md.

For Next.js agents, test whether targeted framework documentation improves results in your own repository. The model covers migrations, data fetching, optimization, caching, and view transitions, and Vercel’s CLI can configure supported agents to use the gateway.

These results use **pass@4**, so a task counts if any of four attempts works rather than measuring first-attempt reliability. ZDR is unavailable, prompts and responses may be used for training, and free access is limited to the stealth period.
Connected Context · Feed7 Judgment

Pixel Canary supplies unusually concrete evidence that targeted repository context can improve agent outcomes: adding Next.js documentation raised pass@4 results from 28 to 30 of 31 tasks. That narrows evaluation toward testing model-plus-context configurations, not model names alone. The small suite, four-attempt scoring, temporary access, and training-data exposure prevent treating 96.8% as first-try reliability or a production default.

Context Map
modelcoding#model-selection#coding-agents#context-engineering
Uncertainty
These results use **pass@4**, so a task counts if any of four attempts works rather than measuring first-attempt reliability. ZDR is unavailable, prompts and responses may be used for training, and free access is limited to the stealth period.