人工智能评估应该与人类合作
arXiv:2608.13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (
共找到 4282 篇相关文章
arXiv:2608.13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (
Reasoning scores keep climbing while per-token compute keeps dropping. GLM-5.2 scores 99.2% on AIME 2026 with about 40 b
【HN用户评论摘要】 To people asking why, this is a good lesson on the Collison’s ambitions. Stripe is one of the best API compan
The Case Against Formal Verification, 50 Years Later Engineers are getting excited about software verification! This may