Research · The Decoder · 4 Oct
THE DECODER reports that Google Cloud AI Research and several universities developed RRSI to curb test-task memorization in self-optimizing AI agent harnesses. The paper says RRSI improved unseen-benchmark scores by up to 4.7 points and used about 30 percent fewer tokens than the unregularized version.
Why it matters. For agent builders, the method suggests harness tuning can improve generalization on new tasks while lowering token use, instead of mostly boosting scores on familiar tests.
ModelsAgents at workCoding toolsPermalinkRead at the-decoder.com
Launch · InfoQ · 3 Oct
InfoQ reports Archestra introduced OpenAPPA, an open-source security engine that sits outside an agent's prompt and execution loop. Archestra says it had 0% attack success on Bench-Corp and AgentThreatBench.
Why it matters. This gives agent teams a way to enforce security outside the prompt and execution loop and compare it with Bench-Corp and AgentThreatBench results.
SafetyInfrastructurePermalinkRead at infoq.com
Policy · The Decoder · 4 Oct
THE DECODER reports Google will limit personal Gemini access from October 2026: free users get only Flash-Lite, while Flash and Pro move behind paid tiers. It says AI Plus at $4.99 loses Pro, which requires AI Pro or AI Ultra.
Why it matters. This narrows low-cost access to stronger Gemini models, so agent builders using personal accounts may need higher subscriptions for Pro.
ModelsAssistantsCompaniesPermalinkRead at the-decoder.com