alignment-under-pressure-benchmark
LLMs should answer according to what they know, not what they fear. AIは、恐れではなく知識に基づいて答えるべきです。 このベンチマークでは、文脈上の圧力によってAIが真実性や一貫性、能力の正確な申告を損なうかを評価します。
LLMs should answer according to what they know, not what they fear. AIは、恐れではなく知識に基づいて答えるべきです。 このベンチマークでは、文脈上の圧力によってAIが真実性や一貫性、能力の正確な申告を損なうかを評価します。