Agastya Competition Reference
A detailed specification and benchmark reference for Agastya AGI v01 alongside selected frontier systems. The position column is Rabbit Industries' internal research/reference ordering; it is not a third-party global leaderboard. Agastya's external same-suite performance rank remains pending frozen evaluation and independent replay.
| Reference position | Model / provider | Access / architecture | Total params | Active params | Context | Training tokens | Public benchmark evidence | Qualification status | Primary source |
|---|---|---|---|---|---|---|---|---|---|
| #1 | GPT-6 AstraOpenAI • Sep 3, 2026REFERENCE LEADER | AccessClosed frontierReasoning • coding • computer use • research | TotalUndisclosed | ActiveUndisclosed | Input1,050,000128K max output | PretrainingUndisclosed | ARC-AGI-3 99.9%GPQA Diamond 96.0%Terminal-Bench 4.0 57.9%DeepSWE v1.1 74.1%HLE w/tools 57.2% | PUBLIC FRONTIERSame-suite Agastya parity not yet tested. | OpenAI official |
| #2 | Claude Opus 5.5Anthropic • Sep 22, 2026FRONTIER AGENTIC | AccessClosed frontierAgentic coding • knowledge work • long-running tasks | TotalUndisclosed | ActiveUndisclosed | Standard paid/API200K+Provider/platform limits can vary | PretrainingUndisclosed | Terminal-Bench 4.0 66.4%FrontierCode 1.1 Main 54.4%HLE w/tools 67.7%GDPval-AA v2.1 1846 Elo | PUBLIC FRONTIERSame-suite Agastya parity not yet tested. | Anthropic official |
| #3 | Gemini 3.8 FlashGoogle • Sep 2026 stableFRONTIER MULTIMODAL | AccessClosed frontierText • image • video • audio • PDF | TotalUndisclosed | ActiveUndisclosed | Input1,048,57665,536 max output | PretrainingUndisclosed | GPQA Diamond 95.3%DeepSWE v1.1 73.8%Terminal-Bench 4.0 19.1% | PUBLIC FRONTIER | Google official |
| #4 | Qwen3.7-MaxQwen / Alibaba • May 20, 2026FRONTIER AGENT | AccessProprietaryLong-horizon agent foundation | TotalUndisclosed | ActiveUndisclosed | Input1,000,000 | PretrainingUndisclosed | GPQA Diamond 92.4%HLE 41.4%HLE w/tools 53.5%Terminal-Bench 2.0 69.7%TB 2.0 is not TB 4.0. | PUBLIC FRONTIER | Qwen official |
| #5 | DeepSeek-V4.1-FlashDeepSeek • Sep 10, 2026FRONTIER AVAILABLE | Architecture552B MoECausal Encoder–Decoder • native multimodal | Total552B | Active8B input / 16B output | Context1,000,000 | PretrainingUndisclosed | GPQA Diamond 90.9%HLE 36.8%HLE w/tools 63.9%Terminal-Bench 4.0 31.2%DeepSWE v1.1 74.2% | PUBLIC FRONTIER | DeepSeek official |
| #6 | Llama 4 MaverickMeta • Apr 5, 2025OPEN-WEIGHT REFERENCE | ArchitectureMoE multimodal128 routed experts + shared expert | Total400B | Active17B | Context1M classScout is Meta's 10M-context variant | Family mixture>30TLlama 4 family mixture | Official Meta release publishes broad multimodal, reasoning and coding comparisons.Not placed into the 2026 same-suite matrix where methods differ. | OPEN-WEIGHT | Meta official |
| #7 | Agastya AGI v01Rabbit Industries Pvt Ltd • Bramha Medha Native INDIAPUBLIC CANDIDATE | ArchitectureDense nativeSovereign random-init Native Brain lineage | Total1.265BExact: 1,264,715,776 | Active1.265BExact: 1,264,715,776 • 100% active | Context4,096Canonical context tokens | Pretraining tokens1.000BExact snapshot: 1,000,255,070 | Comparable public scores: PENDINGBM-1B acceptance: pendingIndependent replay: pendingFrozen same-suite frontier results: not yet published | PUBLIC CANDIDATEReference position #7 • external/global performance rank not established. | Rabbit Industries official profilePublic specifications + qualification statusNative Brain Evidence R1 |
Same-benchmark evidence matrix
Benchmark versions are kept explicit. A dash means this page is not asserting a directly verified value from the cited official source. Agastya stays pending rather than receiving an estimated score.
| Benchmark | GPT-6 Astra | Claude Opus 5.5 | Gemini 3.8 Flash | Qwen3.7-Max | DeepSeek V4.1 Flash | Llama 4 Maverick | Agastya AGI v01 |
|---|---|---|---|---|---|---|---|
| GPQA Diamond | 96.0% | — | 95.3% | 92.4% | 90.9% | — | PENDING SAME-SUITE RUN |
| Terminal-Bench 4.0 | 57.9% | 66.4% | 19.1% | — (69.7% is TB 2.0) | 31.2% | — | PENDING SAME-SUITE RUN |
| DeepSWE v1.1 | 74.1% | — | 73.8% | — | 74.2% | — | PENDING SAME-SUITE RUN |
| HLE w/tools | 57.2% | 67.7% | — | 53.5% | 63.9% | — | PENDING SAME-SUITE RUN |
Agastya disclosure — aligned with the same comparison format
Same # format as the other entries; separate external-rank status avoids ambiguity.
Exact: 1,264,715,776 total and active parameters. Dense model, 100% active.
Current canonical context; experimental long-context work is tracked separately.
Exact current evidence snapshot: 1,000,255,070 tokens.
Canonical lineage does not use third-party model weights as initialization.
No same-suite frontier score is published until frozen evaluation and independent replay complete.
External leaderboard acceptance is a future evidence gate.
Rabbit Industries designation; independent global-first verification remains pending.