Add CHI-Bench eval results — agent harness: OpenAI Agents SDK

#197
Files changed (1) hide show
  1. .eval_results/chi-bench.yaml +40 -0
.eval_results/chi-bench.yaml ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Place at .eval_results/chi-bench.yaml in the DeepSeek V4 Pro model repo.
2
+ # CONFIRM the exact repo id before submitting (org is deepseek-ai).
3
+ # Submit via the model's Community tab as a PR; shows "community-provided" until merged.
4
+ # Values are pass@1 (%) for the best-performing harness for this model: OpenAI Agents SDK.
5
+ - dataset:
6
+ id: actava/chi-bench
7
+ task_id: chi_bench
8
+ value: 14.2
9
+ date: "2026-05-08"
10
+ source:
11
+ url: https://arxiv.org/abs/2605.16679
12
+ name: CHI-Bench
13
+ notes: "Harness: OpenAI Agents SDK; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"
14
+ - dataset:
15
+ id: actava/chi-bench
16
+ task_id: prior_authorization
17
+ value: 10.7
18
+ date: "2026-05-08"
19
+ source:
20
+ url: https://arxiv.org/abs/2605.16679
21
+ name: CHI-Bench
22
+ notes: "Harness: OpenAI Agents SDK; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"
23
+ - dataset:
24
+ id: actava/chi-bench
25
+ task_id: utilization_management
26
+ value: 28.0
27
+ date: "2026-05-08"
28
+ source:
29
+ url: https://arxiv.org/abs/2605.16679
30
+ name: CHI-Bench
31
+ notes: "Harness: OpenAI Agents SDK; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"
32
+ - dataset:
33
+ id: actava/chi-bench
34
+ task_id: care_management
35
+ value: 4.0
36
+ date: "2026-05-08"
37
+ source:
38
+ url: https://arxiv.org/abs/2605.16679
39
+ name: CHI-Bench
40
+ notes: "Harness: OpenAI Agents SDK; Protocol: 75 tasks x 3 trials; Metric: pass@1 (%)"