AI快讯 / AI 开源项目

codejudge-ai

Deterministic evaluation and reproducible benchmarking for AI-generated code with hardened Docker sandboxing, trusted tests, static analysis, and multi-provider model comparison.

原文来源AI 开源 Releases
查看官方原文

Deterministic evaluation and reproducible benchmarking for AI-generated code with hardened Docker sandboxing, trusted tests, static analysis, and multi-provider model comparison.