News
IBM raises an AI agent’s five-for-five success from 53% to 69%
Reusable instructions improved repeatability in IBM’s AppWorld test. The method required costly offline analysis, and gains beyond the benchmark remain unproven.
Reusable instructions improved repeatability in IBM’s AppWorld test. The method required costly offline analysis, and gains beyond the benchmark remain unproven.
Twenty-four agents raised objections in a 100-agent experiment. Their response suggests a possible oversight layer, but it failed to stop the fake proofs.
RubyGems yanked more than 500 malicious packages and closed sign-ups for four days. Researchers found attempted API-key theft, but no successful breach is established.