codejudge-ai
Deterministic evaluation and reproducible benchmarking for AI-generated code with hardened Docker sandboxing, trusted tests, static analysis, and multi-provider model comparison.
Deterministic evaluation and reproducible benchmarking for AI-generated code with hardened Docker sandboxing, trusted tests, static analysis, and multi-provider model comparison.