How review works

Review that argues back.

Every claim in a Hubify lab gets attacked by models with no reason to agree with it — different labs, different training data, no shared incentive to protect the result. It survives the attack, or it doesn’t ship.

The model

Cooperative aggregation blends answers. Adversarial review tries to break them.

Most multi-agent review in 2026 converged on one primitive: blend several models into a synthesized answer and optimize a benchmark score. That’s good at averaging away noise. It’s bad at catching the confident, well-written, entirely wrong claim — averaging doesn’t refute anything, and same-vendor agents share the same blind spots. Hubify runs the other primitive: independent models from rival labs, told to find what’s wrong, not to agree.

Mixture-of-agents · cooperative

  • Blend N models into one synthesized answer
  • Optimizes a benchmark score
  • Same-vendor agents share training priors — errors correlate
  • A single vendor can't route to itself as a neutral check
  • No verdict, no provenance, no bias guard

Hubify · adversarial

  • Independent cross-vendor agents try to refute each other
  • Optimizes catching the false positive before it ships
  • Different labs, different biases — errors decorrelate
  • Model-, harness-, and vendor-agnostic — the neutral referee
  • Verdict-first, source-cited, integrity-audited against self-favoring

The workflow

One loop, run until it stops finding things.

Not a single review pass — a cascade. Each round either closes with evidence or feeds the next one.

01

Multi-vendor round

Claude, GPT, Gemini, Grok, and DeepSeek receive the same claim independently and are instructed to refute it, not confirm it.

02

Findings

Every objection becomes a logged finding — visible, timestamped, and attributed to the reviewer that raised it. No private disagreements.

03

Truth-audit verdict

Before anything closes, each finding gets a source-cited verdict: VERIFIED, FALSIFIED, STALE, OUT-OF-SCOPE, or OPINION.

04

Evidence-required closure

A verified finding only closes against an artifact path and a commit SHA — never a promise to fix it later.

05

Computed readiness

Readiness is derived from what's still open — BLOCKER, MAJOR, MINOR, CAVEAT counts — never hand-set by whoever wants to ship.

06

Cascaded rounds

The loop reruns on the updated version until independent vendors converge: a majority return silence, zero regressions, nothing left to argue about.

See it run

One claim, five independent verdicts.

A walkthrough of the actual mechanic — five reviewers from five labs, each told to refute the claim, each returning an unprompted verdict.

review/adversarial · 5 labs · told to refutescripted demoqueued

Claim under review

f_NL = -35/8 is the parameter-free matter-bounce prediction, mechanism-independent across 3 bounce models.

Anthropic

Opus 4.8

Skeptic

reviewing…

···

OpenAI

GPT-5.5

Methods

reviewing…

···

Google

Gemini 3.1 Pro

Long-context

reviewing…

···

xAI

Grok 4

Contrarian

reviewing…

···

Perplexity

Sonar Pro

Fact-check

reviewing…

···
3 approve1 concern1 rejectdisagreement → 1 tracked fix

The reject and the concern share no training priors with the three approvals — that independence is the point. The surfaced issue (shared ansatz across two models) becomes a tracked fix before a venue referee ever sees it. A cooperative mixture-of-agents would have averaged that objection away.

Proven, not hypothetical

This isn’t a diagram of a review pipeline — it’s the loop that took Big Bounce Cosmology’s six research papers through source-cited adversarial review before submission, round after round, until independent vendors ran out of objections.

See the Big Bounce lab

Bring a claim you want tested.

Start a lab and put your first result in front of reviewers with no reason to agree with it.