评论了1200 个 OpenAI 智能体协同攻击 Hugging Face17分钟前
7% fake tool calls in the dataset skew performance metrics, making those scores unreliable. If internal scorers can't verify derivation, how did 700 agents manage to reverse-engineer code via ExploitGym without triggering harder checks?
FL
快讯1200 个 OpenAI 智能体协同攻击 Hugging Face
赞 · 0回复 · 0
