# Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Source: [arXiv](https://arxiv.org/abs/2609.15983v1)  
Feed7 permalink: https://feed7.dev/p/2609-15983v1-13dydbk  
Published: 2026-09-14T17:58:55.000Z  
Trust: Needs Review (needs_review)

## Why Included

Stellar Colosseum coordinates parallel strategy search, falsification, decomposition, and verifier feedback for long research tasks. Its harness patterns may transfer to agents handling interdependent coding work.

## Source Summary

Stellar Colosseum is a model-agnostic many-agent harness for long-horizon mathematics and theoretical computer science. It explores competing strategies, gates decomposition on readiness, attacks candidates through falsification, and routes verifier findings back to affected proof sections.

## Practical Implication

Agent builders should reconsider immediate task decomposition: search and challenge candidate plans first, then split only mature routes into linked subproblems. The same workflow appears in Google Antigravity Teamwork as the **Long Proof pattern**.

## Agent-Ready Context

Stellar Colosseum is a model-agnostic many-agent harness for long-horizon mathematics and theoretical computer science. It explores competing strategies, gates decomposition on readiness, attacks candidates through falsification, and routes verifier findings back to affected proof sections.

Agent builders should reconsider immediate task decomposition: search and challenge candidate plans first, then split only mature routes into linked subproblems. The same workflow appears in Google Antigravity Teamwork as the **Long Proof pattern**.

With specified Gemini models, the harness reports **71.0% on TCS-Bench** and **218 of 222 Codeforces problems** solved with execution feedback. Those results concern proofs and competitive programming, so transfer to general software projects remains untested here.

## Connected Context

Feed7 judgment across 778 accumulated Signals:

Stellar Colosseum makes delayed decomposition a concrete many-agent control pattern: explore and falsify candidate routes before splitting mature ones, then send verifier findings back to the affected proof sections. This sharpens prior long-horizon guidance around compositional trajectories and evidence-backed completion, while its strong reported results remain bounded to proofs and competitive programming.

- [The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation](https://feed7.dev/p/2607-24720v1-0gihy13) — The controlled planning evidence says compositional trajectories and explicit transitions matter; Stellar Colosseum supplies a harness consequence by postponing decomposition until a route is ready and preserving links among subproblems.
- [What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip](https://feed7.dev/p/what-does-done-even-mean-agents-and-paperclip-s-liveness-model-dotta-pap-0lx8wfc) — The liveness model’s evidence-based definition of done complements the harness’s falsification and verifier-feedback loops, preventing an agent’s claimed proof completion from serving as final approval.
- [Don't Build Agents You Can't Answer For — Addy Osmani](https://feed7.dev/p/don-t-build-agents-you-can-t-answer-for-addy-osmani-1y9nwej) — Stellar Colosseum operationalizes the demand for inspectable evidence by routing verifier findings to specific proof sections, though its reported domains do not establish the same accountability for software changes.
- [alphaXiv/OpenResearch](https://feed7.dev/p/openresearch-1472dzs) — OpenResearch contributes isolated execution and experiment lineage for parallel research, while Stellar Colosseum contributes strategy competition, readiness-gated decomposition, and targeted falsification; together they cover complementary research-harness controls.

## Context Map

- Layer: agent
- Domains: coding, research
- Topics: multi-agent, harness-engineering, agent-reliability

## Uncertainty

- With specified Gemini models, the harness reports **71.0% on TCS-Bench** and **218 of 222 Codeforces problems** solved with execution feedback. Those results concern proofs and competitive programming, so transfer to general software projects remains untested here.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
