Benchmark degli Agent
Benchmark degli agenti di τ³-Banking Knowledge
OpenGPT mantiene 21 righe pubblicate associate alla versione esatta del benchmark τ³-Banking Knowledge e all’ambiente mostrato sotto.
- Attività
- 97
- Righe pubblicate
- 21
- Data dei dati della fonte
- Identità esaminata
- Attuale
Ambito delle evidenze
v1.0.1 · AllTools · GPT-5.2 low user simulator · seed 300 · 4 trials. banking knowledge tasks. Standard scaffold with BM25, embedding retrieval, and sandboxed command-line search. 4 trials per task · 388 trajectories per configuration.
How OpenGPT handles this evidence
OpenGPT conserva la metrica, l’ambito delle attività, l’ambiente, la versione, i campi della configurazione e la data della fonte indicati dal publisher. Non crea un punteggio complessivo degli Agent tra benchmark diversi.
Limiti del confronto
- Un risultato più elevato si applica solo alla versione del benchmark e alle condizioni visualizzate.
- I risultati degli Agent riflettono il sistema completo, non soltanto il modello di base.
- I dati mancanti su costi, affidabilità, sicurezza o latenza restano non indicati anziché essere dedotti.
- Le righe incluse utilizzano la stessa versione del benchmark, la suddivisione Banking Knowledge di 97 attività, la configurazione AllTools, il simulatore utente, il seed e il protocollo di quattro prove. I costi mancanti rimangono non dichiarati. Questa istantanea include tutti i 21 invii compatibili degli editori disponibili nella revisione bloccata.
v1.0.1 · AllTools · GPT-5.2 low user simulator · seed 300 · 4 trials
Visualizzate 21 di 21 righe pubblicate. Ogni risultato rimane associato al benchmark e alla configurazione della fonte.
- τ³ Banking AllTools · Qwen 3.8 Max · xhighQwen · Qwen 3.8 Max#1Pass¹: 55.15%τ³ standard agent
- τ³ Banking AllTools · Claude Opus 5 · maxAnthropic · Claude Opus 5#2Pass¹: 48.71%τ³ standard agent
- τ³ Banking AllTools · Grok 4.5 · altoxAI · Grok 4.5#3Pass¹: 47.94%τ³ standard agent
- τ³ Banking AllTools · GPT-5.6-sol · xhighOpenAI · GPT-5.6-sol#4Pass¹: 46.91%τ³ standard agent
- τ³ Banking AllTools · GPT-5.5 · xhighOpenAI · GPT-5.5#5Pass¹: 44.59%τ³ standard agent
- τ³ Banking AllTools · Muse Spark 1.1 · xhighMeta · Muse Spark 1.1#6Pass¹: 40.46%τ³ standard agent
- τ³ Banking AllTools · Claude Opus 4.7 · maxAnthropic · Claude Opus 4.7#7Pass¹: 40.21%τ³ standard agent
- τ³ Banking AllTools · Claude Fable 5 · maxAnthropic · Claude Fable 5#8Pass¹: 39.69%τ³ standard agent
- τ³ Banking AllTools · Claude Opus 4.8 · maxAnthropic · Claude Opus 4.8#9Pass¹: 39.69%τ³ standard agent
- τ³ Banking AllTools · GPT-5.4 · xhighOpenAI · GPT-5.4#10Pass¹: 39.43%τ³ standard agent
- τ³ Banking AllTools · GLM-5.2 · xhighZ.ai · GLM-5.2#11Pass¹: 37.11%τ³ standard agent
- τ³ Banking AllTools · Kimi K3 · maxMoonshot AI · Kimi K3#12Pass¹: 37.11%τ³ standard agent
- τ³ Banking AllTools · GPT-5.2 · altoOpenAI · GPT-5.2#13Pass¹: 32.22%τ³ standard agent
- τ³ Banking AllTools · Claude Opus 4.6 · maxAnthropic · Claude Opus 4.6#14Pass¹: 27.32%τ³ standard agent
- τ³ Banking AllTools · Gemini 3.1 Pro Preview · altoGoogle · Gemini 3.1 Pro Preview#15Pass¹: 26.03%τ³ standard agent
- τ³ Banking AllTools · Inkling · maxThinking Machines Lab · Inkling#16Pass¹: 25.00%τ³ standard agent
- τ³ Banking AllTools · Claude Opus 4.5 · altoAnthropic · Claude Opus 4.5#17Pass¹: 24.74%τ³ standard agent
- τ³ Banking AllTools · Grok 4.2 · altoxAI · Grok 4.2#18Pass¹: 18.04%τ³ standard agent
- τ³ Banking AllTools · Grok 4 fast · altoxAI · Grok 4 fast#19Pass¹: 15.72%τ³ standard agent
- τ³ Banking AllTools · Gemini 2.5 Pro · altoGoogle · Gemini 2.5 Pro#20Pass¹: 13.66%τ³ standard agent
- τ³ Banking AllTools · Grok 4.1 fast · altoxAI · Grok 4.1 fast#21Pass¹: 13.14%τ³ standard agent