⚡ September 14, 2026¶
Generated: 2026-09-14 20:04 UTC
Total Duration: 1h 33m 59s
Iterations: 1
Judge (classifier) model: gpt-4.1
Fast Benchmark
Markers: regression or benchmark
Schedule: Weekly (Sunday 2 AM UTC)
Purpose: Quick regression tests to catch breaking changes
HolmesGPT is continuously evaluated against real-world Kubernetes and cloud troubleshooting scenarios.
If you find scenarios that HolmesGPT does not perform well on, please consider adding them as evals to the benchmark.
Model Accuracy Comparison¶
| Model | Pass | Fail | Skip/Error | Total | Success Rate |
|---|---|---|---|---|---|
| glm-5.3 | 56 | 7 | 0 | 63 | 🟡 89% (56/63) |
| haiku-4.5 | 51 | 12 | 0 | 63 | 🟡 81% (51/63) |
| kimi-k3 | 57 | 6 | 0 | 63 | 🟡 90% (57/63) |
| opus-5 | 58 | 5 | 0 | 63 | 🟡 92% (58/63) |
| sonnet-5 | 52 | 11 | 0 | 63 | 🟡 83% (52/63) |
Model Cost Comparison¶
| Model | Tests | Avg Cost | Min Cost | Max Cost | Total Cost |
|---|---|---|---|---|---|
| glm-5.3 | 63 | $0.09 | $0.00 | $1.92 | $5.60 |
| haiku-4.5 | 63 | $0.04 | $0.02 | $0.12 | $2.57 |
| kimi-k3 | 63 | $0.16 | $0.00 | $3.24 | $9.93 |
| opus-5 | 63 | $0.36 | $0.10 | $1.19 | $22.59 |
| sonnet-5 | 63 | $0.11 | $0.04 | $0.38 | $7.08 |
Model Latency Comparison¶
| Model | Avg (s) | Min (s) | Max (s) | P50 (s) | P95 (s) |
|---|---|---|---|---|---|
| glm-5.3 | 64.0 | 9.7 | 734.8 | 29.7 | 148.6 |
| haiku-4.5 | 28.6 | 11.7 | 61.1 | 24.8 | 54.9 |
| kimi-k3 | 120.5 | 10.1 | 1578.9 | 51.9 | 365.0 |
| opus-5 | 71.6 | 11.2 | 310.3 | 51.7 | 222.7 |
| sonnet-5 | 49.1 | 11.0 | 151.7 | 35.2 | 131.0 |
Performance by Tag¶
Success rate by test category and model:
| Tag | glm-5.3 | haiku-4.5 | kimi-k3 | opus-5 | sonnet-5 | Warnings |
|---|---|---|---|---|---|---|
| benchmark | 🟡 78% (14/18) | 🟡 72% (13/18) | 🟡 83% (15/18) | 🟡 89% (16/18) | 🟡 72% (13/18) | |
| context_window | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟡 83% (⅚) | |
| counting | 🟢 100% (6/6) | 🟡 83% (⅚) | 🟢 100% (6/6) | 🟡 67% (4/6) | 🟢 100% (6/6) | |
| datetime | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟡 89% (8/9) | |
| easy | 🟡 88% (21/24) | 🟡 92% (22/24) | 🟡 88% (21/24) | 🟡 96% (23/24) | 🟡 79% (19/24) | |
| elasticsearch | 🟢 100% (9/9) | 🟡 56% (5/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | |
| grafana | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | |
| hard | 🟡 67% (4/6) | 🟡 50% (3/6) | 🟡 83% (⅚) | 🟢 100% (6/6) | 🟡 67% (4/6) | |
| kubernetes | 🟡 90% (27/30) | 🟡 87% (26/30) | 🟡 83% (25/30) | 🟡 83% (25/30) | 🟡 77% (23/30) | |
| logs | 🟡 72% (13/18) | 🟡 61% (11/18) | 🟡 72% (13/18) | 🟡 83% (15/18) | 🟡 56% (10/18) | |
| loki | 🟡 50% (3/6) | 🟡 50% (3/6) | 🟡 33% (2/6) | 🟡 50% (3/6) | 🟡 33% (2/6) | |
| medium | 🟡 93% (25/27) | 🟡 78% (21/27) | 🟡 93% (25/27) | 🟡 93% (25/27) | 🟡 85% (23/27) | |
| metrics | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | |
| multi-cluster | 🟢 100% (9/9) | 🟡 56% (5/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | |
| network | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟡 67% (⅔) | 🟢 100% (3/3) | 🟡 33% (⅓) | |
| one-test | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | |
| port-forward | 🟡 67% (6/9) | 🟡 67% (6/9) | 🟡 56% (5/9) | 🟡 67% (6/9) | 🟡 56% (5/9) | |
| question-answer | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | |
| regression | 🟡 93% (42/45) | 🟡 84% (38/45) | 🟡 93% (42/45) | 🟡 93% (42/45) | 🟡 87% (39/45) | |
| skills | 🟡 92% (33/36) | 🟡 92% (33/36) | 🟡 92% (33/36) | 🟡 92% (33/36) | 🟡 83% (30/36) | |
| transparency | 🟢 100% (9/9) | 🟡 56% (5/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | |
| Overall | 🟡 89% (56/63) | 🟡 81% (51/63) | 🟡 90% (57/63) | 🟡 92% (58/63) | 🟡 83% (52/63) |
Raw Results¶
Status of all evaluations across models. Color coding:
- 🟢 Passing 100% (stable)
- 🟡 Passing 1-99%
- 🔴 Passing 0% (failing)
- 🔧 Mock data failure (missing or invalid test data)
- ⚠️ Setup failure (environment/infrastructure issue)
- ⏱️ Timeout or rate limit error
- ⏭️ Test skipped (e.g., known issue or precondition not met)
| Eval ID | glm-5.3 | haiku-4.5 | kimi-k3 | opus-5 | sonnet-5 |
|---|---|---|---|---|---|
| 09_crashpod 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 100a_loki_historical_logs 🔗 | 🟡 | 🟡 | 🟡 | 🟡 | 🟡 |
| 101_loki_historical_logs_pod_deleted 🔗 | 🟡 | 🟡 | 🟡 | 🟡 | 🟡 |
| 108_logs_nearby_lines 🔗 | 🟡 | 🔴 | 🟡 | 🟢 | 🟡 |
| 112_find_pvcs_by_uuid 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 12_job_crashing 🔗 | 🟡 | 🟢 | 🟢 | 🟢 | 🟢 |
| 176_network_policy_blocking_traffic_no_skills 🔗 | 🟢 | 🟢 | 🟡 | 🟢 | 🟡 |
| 179_grafana_big_dashboard_query 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 227_count_configmaps_per_namespace[0] 🔗 | 🟢 | 🟡 | 🟢 | 🟡 | 🟢 |
| 243_pod_names_contain_service 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟡 |
| 24_misconfigured_pvc 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 254_elasticsearch_dr_test_log_check 🔗 | 🟢 | 🔴 | 🟢 | 🟢 | 🟢 |
| 259_wrong_cluster_logs_confusion 🔗 | 🟢 | 🟡 | 🟢 | 🟢 | 🟢 |
| 260_global_es_remote_cluster_logs 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 283_todowrite_multistep_audit 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 43_current_datetime_from_prompt 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 51_logs_summarize_errors 🔗 | 🟢 | 🟡 | 🟢 | 🟢 | 🟡 |
| 61_exact_match_counting 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 73a_time_window_anomaly 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| 73b_time_window_anomaly 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟡 |
| 96_no_matching_skill 🔗 | 🟢 | 🟢 | 🟢 | 🟢 | 🟢 |
| SUMMARY | 🟡 89% (56/63) | 🟡 81% (51/63) | 🟡 90% (57/63) | 🟡 92% (58/63) | 🟡 83% (52/63) |
Detailed Raw Results¶
| Eval ID | glm-5.3 | haiku-4.5 | kimi-k3 | opus-5 | sonnet-5 |
|---|---|---|---|---|---|
| 09_crashpod 🔗 | 🟢 100% (3/3) / ⏱️ 17.6s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 22.7s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 54.8s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 36.0s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 28.2s / 💰 $0.07 |
| 100a_loki_historical_logs 🔗 | 🟡 33% (⅓) / ⏱️ 266.0s / 💰 $0.22 | 🟡 33% (⅓) / ⏱️ 33.1s / 💰 $0.05 | 🟡 33% (⅓) / ⏱️ 278.3s / 💰 $0.26 | 🟡 33% (⅓) / ⏱️ 195.8s / 💰 $0.86 | 🟡 33% (⅓) / ⏱️ 113.8s / 💰 $0.21 |
| 101_loki_historical_logs_pod_deleted 🔗 | 🟡 67% (⅔) / ⏱️ 466.3s / 💰 $0.93 | 🟡 67% (⅔) / ⏱️ 25.1s / 💰 $0.03 | 🟡 33% (⅓) / ⏱️ 198.6s / 💰 $0.27 | 🟡 67% (⅔) / ⏱️ 274.9s / 💰 $0.91 | 🟡 33% (⅓) / ⏱️ 84.6s / 💰 $0.14 |
| 108_logs_nearby_lines 🔗 | 🟡 33% (⅓) / ⏱️ 36.3s / 💰 $0.04 | 🔴 0% (0/3) / ⏱️ 39.2s / 💰 $0.06 | 🟡 67% (⅔) / ⏱️ 40.9s / 💰 $0.09 | 🟢 100% (3/3) / ⏱️ 59.8s / 💰 $0.34 | 🟡 33% (⅓) / ⏱️ 51.0s / 💰 $0.13 |
| 112_find_pvcs_by_uuid 🔗 | 🟢 100% (3/3) / ⏱️ 18.0s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 18.9s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 26.7s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 35.3s / 💰 $0.21 | 🟢 100% (3/3) / ⏱️ 22.2s / 💰 $0.06 |
| 12_job_crashing 🔗 | 🟡 33% (⅓) / ⏱️ 31.3s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 20.8s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 29.5s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 32.2s / 💰 $0.20 | 🟢 100% (3/3) / ⏱️ 33.5s / 💰 $0.08 |
| 176_network_policy_blocking_traffic_no_skills 🔗 | 🟢 100% (3/3) / ⏱️ 27.6s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 36.9s / 💰 $0.06 | 🟡 67% (⅔) / ⏱️ 83.0s / 💰 $0.13 | 🟢 100% (3/3) / ⏱️ 55.2s / 💰 $0.30 | 🟡 33% (⅓) / ⏱️ 101.2s / 💰 $0.27 |
| 179_grafana_big_dashboard_query 🔗 | 🟢 100% (3/3) / ⏱️ 28.7s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 19.1s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 24.2s / 💰 $0.05 | 🟢 100% (3/3) / ⏱️ 34.2s / 💰 $0.25 | 🟢 100% (3/3) / ⏱️ 24.4s / 💰 $0.08 |
| 227_count_configmaps_per_namespace[0] 🔗 | 🟢 100% (3/3) / ⏱️ 25.4s / 💰 $0.05 | 🟡 67% (⅔) / ⏱️ 28.0s / 💰 $0.06 | 🟢 100% (3/3) / ⏱️ 41.0s / 💰 $0.04 | 🟡 33% (⅓) / ⏱️ 48.8s / 💰 $0.30 | 🟢 100% (3/3) / ⏱️ 30.4s / 💰 $0.14 |
| 243_pod_names_contain_service 🔗 | 🟢 100% (3/3) / ⏱️ 25.8s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 23.6s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 42.5s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 36.6s / 💰 $0.20 | 🟡 67% (⅔) / ⏱️ 34.2s / 💰 $0.07 |
| 24_misconfigured_pvc 🔗 | 🟢 100% (3/3) / ⏱️ 28.9s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 24.7s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 29.5s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 36.4s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 31.8s / 💰 $0.07 |
| 254_elasticsearch_dr_test_log_check 🔗 | 🟢 100% (3/3) / ⏱️ 52.7s / 💰 $0.05 | 🔴 0% (0/3) / ⏱️ 49.7s / 💰 $0.06 | 🟢 100% (3/3) / ⏱️ 263.8s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 85.9s / 💰 $0.35 | 🟢 100% (3/3) / ⏱️ 53.8s / 💰 $0.11 |
| 259_wrong_cluster_logs_confusion 🔗 | 🟢 100% (3/3) / ⏱️ 61.3s / 💰 $0.10 | 🟡 67% (⅔) / ⏱️ 46.5s / 💰 $0.07 | 🟢 100% (3/3) / ⏱️ 450.2s / 💰 $0.66 | 🟢 100% (3/3) / ⏱️ 107.2s / 💰 $0.56 | 🟢 100% (3/3) / ⏱️ 54.1s / 💰 $0.11 |
| 260_global_es_remote_cluster_logs 🔗 | 🟢 100% (3/3) / ⏱️ 35.3s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 54.6s / 💰 $0.07 | 🟢 100% (3/3) / ⏱️ 607.1s / 💰 $1.15 | 🟢 100% (3/3) / ⏱️ 81.7s / 💰 $0.33 | 🟢 100% (3/3) / ⏱️ 75.5s / 💰 $0.13 |
| 283_todowrite_multistep_audit 🔗 | 🟢 100% (3/3) / ⏱️ 32.4s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 24.9s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 35.0s / 💰 $0.06 | 🟢 100% (3/3) / ⏱️ 38.6s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 33.7s / 💰 $0.08 |
| 43_current_datetime_from_prompt 🔗 | 🟢 100% (3/3) / ⏱️ 10.0s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 12.7s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 10.3s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 11.9s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 11.4s / 💰 $0.04 |
| 51_logs_summarize_errors 🔗 | 🟢 100% (3/3) / ⏱️ 19.0s / 💰 $0.02 | 🟡 67% (⅔) / ⏱️ 18.6s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 62.3s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 33.9s / 💰 $0.23 | 🟡 67% (⅔) / ⏱️ 28.5s / 💰 $0.07 |
| 61_exact_match_counting 🔗 | 🟢 100% (3/3) / ⏱️ 12.1s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 15.7s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 18.8s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 17.0s / 💰 $0.13 | 🟢 100% (3/3) / ⏱️ 15.2s / 💰 $0.05 |
| 73a_time_window_anomaly 🔗 | 🟢 100% (3/3) / ⏱️ 31.2s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 26.1s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 97.6s / 💰 $0.09 | 🟢 100% (3/3) / ⏱️ 102.1s / 💰 $0.61 | 🟢 100% (3/3) / ⏱️ 54.2s / 💰 $0.14 |
| 73b_time_window_anomaly 🔗 | 🟢 100% (3/3) / ⏱️ 64.5s / 💰 $0.05 | 🟢 100% (3/3) / ⏱️ 25.0s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 56.4s / 💰 $0.07 | 🟢 100% (3/3) / ⏱️ 112.7s / 💰 $0.59 | 🟡 67% (⅔) / ⏱️ 55.6s / 💰 $0.13 |
| 96_no_matching_skill 🔗 | 🟢 100% (3/3) / ⏱️ 53.8s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 34.9s / 💰 $0.05 | 🟢 100% (3/3) / ⏱️ 80.6s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 66.3s / 💰 $0.40 | 🟢 100% (3/3) / ⏱️ 94.0s / 💰 $0.19 |
Results are automatically generated and updated weekly. View full traces and detailed analysis in Braintrust experiment: ci-benchmark-34880758879.