24小时 AI快讯 实时更新
聚合全球 AI 厂商、模型、API、开源项目与服务状态动态。
LATEST UPDATES
最新获取
2026-09-03
-
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
arXiv:2609.01611v1 Announce Type: new Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks. EvalDetectBench ships with a newly curated transcript suite covering current frontier system-card evaluations and diverse deployment sources. The benchmark serves two purposes: measuring how reliably frontier LLMs recognize that they are being evaluated, and assessing how detectable individual benchmarks are as evaluations. We identify two methodological choices in the existing literature that introduce systematic bias: the identity of the model that generated the deployment transcripts accounts for 11.25% of measurement variance and can reorder model rankings; and elicitation prompts selected for high performance on one model can perform near chance on others. EvalDetectBench corrects for both via per-model probe calibration and a stratified generator-harmonisation procedure.
arXiv AI 研究研究动态查看详情 -
Sand.ai:全球首个千亿级开源MoE视频生成模型MAGI Preview
Sand.ai开源千亿MoE视频模型,成本大降破局行业难题
行业媒体行业动态查看详情
-
花 5 万澳元上学,却和 AI 聊天机器人上课?澳洲大学改革引发激烈争议
澳洲高校引入AI替代面授,引发师生争议。
行业媒体行业动态查看详情
-
Anthropic 升级 Claude Code / Cowork,可后台操控用户 Mac
Anthropic 宣布升级 Claude 的电脑控制功能,现可在 macOS 后台运行,同步处理点击、输入等操作,不影响用户其他工作。该功能支持 Cowork 和 Code 产品,仅限 Pro 和 Max 订阅用户使用。#Claude# #AI助手#
行业媒体行业动态查看详情
-
Gemini 3.8 Flash上线:性能接近 Opus 5,价格不变但任务更贵
6周连更3次,留给Gemini 3的版本号不多了。
行业媒体行业动态查看详情
-
Claude把“回忆”砍到2毛5,OpenAI的万亿参数正在失去护城河
一场关于AI"记性"的定价战,暴露了Agent商业化的真正底牌。
行业媒体行业动态查看详情
-
最强模型Fable 5.1发布,Anthropic要用降价换IPO
最强模型不等于最好卖。
行业媒体行业动态查看详情
-
一个模型场景通吃!它石智航AWE3.7的泛化能力有点狠
从工厂一路干到家庭
行业媒体行业动态查看详情
-
DistributedAI/102-v5
Hugging Face模型更新查看详情 -
Workflows API Degraded
Status: Resolved The incident has been resolved Affected services Workflows API
Mistral服务状态查看详情 -
AI Registry Prompts API Degraded
Status: Resolved The incident has been resolved Affected services Prompts API
Mistral服务状态查看详情 -
sarahjonesbeck/multitask-kaggle52
Hugging Face模型更新查看详情 -
t2ance/atlas-ppo5-process_originalweights_actorlr7p5e6_from_s10
Hugging Face模型更新查看详情 -
Stage-org/appworld-qwen35-4b-luna-no-mcqa-lora-epoch1-iter1
Hugging Face模型更新查看详情 -
t2ance/atlas-ppo5-process_final075_discovery025_actorlr7p5e6_from_s10
Hugging Face模型更新查看详情 -
SilentByte-62p/co3-b
Hugging Face模型更新查看详情 -
sarahjonesbeck/matching
Hugging Face模型更新查看详情 -
DistributedAI/102-v8
Hugging Face模型更新查看详情 -
david425/271a0398-547d-4730-adb5-85ee222b1fa3
Hugging Face模型更新查看详情 -
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
Hugging Face模型更新查看详情