Add community evaluation results for SWE-BENCH_VERIFIED, SWE-BENCH_PRO

#11
by nielsr HF Staff - opened

This PR adds community-provided evaluation results for the following benchmarks:

These results were extracted from the model card. This is based on the new evaluation results feature.

Note: This is an automated PR. Please review the evaluation results before merging.

Kwaipilot org

Thanks for this automated PR.
After consideration, we prefer not to merge it for now. There exist significant discrepancies across different SWE-Bench evaluation setups, and merging results obtained under inconsistent evaluation standards will limit the comparability of our benchmark table.

Hi, thanks.

Note that this feature is based on trust by the community, by extracting reported evals from papers, model cards, blog posts, Github READMEs and more to consolidate the numbers.

Kwaipilot org

@nielsr We recommend removing our model from the ranking list, as otherwise misunderstandings may arise where the outside world assumes our scores are relatively low. Thank you.

Ok, I'll close this PR.

nielsr changed pull request status to closed

Sign up or log in to comment