Peer-reviewed research, distilled weekly

Applied AI Research

Every edition breaks down a peer-reviewed paper into actionable strategy. No hype — just evidence from the world's leading labs.

Previous Editions

hallucination-reductionmulti-agentbenchmarking

The Leaderboard Winner Is the Wrong Model to Deploy

MAS-HQ protocol reveals that systems winning factuality leaderboards often lose when API costs are counted, showing 4× token waste for marginal gains.

Shanghai Jiao Tong University

August 5, 20268 minRead
adversarial-audioprompt-injectionaudio-language-models

The Attack That Hides Instructions Inside Audio Reverberation

Adversarial audio files can hijack AI voice assistants to execute unauthorized actions like downloading malware or exfiltrating user data with 79-96% success rates while sounding normal to users.

Zhejiang University, State Key Laboratory of Blockchain and Data Security, Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Nanyang Technological University, National University of Singapore

August 5, 20269 minRead
voice-interfacehealthcare-monitoringpreventive-care

Voice Agents Closing the Monitoring Gap No Physician Budget Ever Will

Agent PULSE uses voice-based AI to monitor chronic disease patients at scale, achieving 70% acceptance in a 33-patient pilot while reducing costs for routine healthcare tasks.

IBM T.J. Watson Research Center, Nova School of Business and Economics, Morehouse School of Medicine, Cleveland Clinic Foundation

August 5, 20268 minRead
relation-extractionsmall-language-modelsdomain-adaptive-training

A 0.5 GB Model that Outperformed GPT-5.4

Sub-billion parameter models fine-tuned with task-specific data outperform GPT-5.4 and Claude Sonnet on relation extraction by 15-26 F1 points across general and literary benchmarks.

Aristotle University of Thessaloniki, Athena Research Center

July 26, 20269 minRead
retrieval-augmented-generationsmall-language-modelson-device-inference

A 4B Model Matching GPT-5-mini running Locally

Small 4-billion parameter models running on CPU-only hardware achieve quality matching GPT-5-mini in RAG systems, eliminating GPU requirements for enterprise deployment.

Siberian Neuronets LLC

July 19, 20268 minRead
neuro-symbolic-aiknowledge-graph-groundingsmall-language-models

The AI Reasoner That Breaks When It Reads Its Own Notes

Small language models gain 1.5-2x accuracy on kinship reasoning by calling specialized extraction and graph neural network tools, but fail when self-extracting noisy facts.

Institute of Informatics and Telecommunications, National Center for Scientific Research 'Demokritos'

July 19, 20269 minRead
graph-ragmulti-agent-systemsknowledge-graph-construction

The GraphRAG Paradox

Multi-agent system with shared memory builds knowledge graphs that reduce retrieval noise by 50% while achieving 0.061-second query latency.

Xiamen University, Jilin University

July 12, 20268 minRead
graph-augmented-retrievalstructural-query-processingtool-use

The Answer Lives in the Graph

An LLM given nine typed graph primitives as tools outperforms hand-coded query handlers, proving the barrier is operator vocabulary not model intelligence.

Siemens Digital Industries Software

July 5, 20269 minRead
adversarial-attacksjailbreakingllm-security

How Two Blocked Attacks Became One That Wasn't

Hybrid jailbreak methods combining token and prompt attacks bypass defenses on Vicuna-7B with 37-58% success rates where single attacks failed completely.

Purdue University

June 28, 20269 minRead
agentic-architecturetoken-optimizationhallucination-prevention

What If Querying 16 Million Records Cost the Same as 50

RES architecture cuts AI agent token costs to constant 1,574 tokens regardless of dataset size by processing data through deterministic code instead of LLM context windows.

Walmart Tech

June 14, 20269 minRead
compound-ai-systemsmulti-agent-orchestrationenterprise-ai

What Happens When You Stop Building Around the LLM

A blueprint architecture orchestrates enterprise AI agents through streams and registries, enabling multi-modal data access and cost-optimized compound AI workflows.

Megagon Labs

June 6, 20269 minRead
skill-optimizationtext-space-learningagent-adaptation

SkillOpt: a Mechanism That Makes Skills Training Stable

SkillOpt optimizes agent skill documents like neural weights with validation gates and edit budgets, improving task accuracy by up to 39 points without model retraining.

Microsoft, Shanghai Jiao Tong University, Tongji University, Fudan University

May 28, 20268 minRead
agent-governanceai-safetyalignment

Safe Alone, Dangerous Together: The AI Agent Blind Spot

A governance taxonomy organizes AI agent interventions into five categories—alignment, control, visibility, security, and societal integration—to manage risks as agents approach human-level task performance.

IAPS (Institute for AI Policy and Strategy)

May 16, 20269 minRead
retrieval-augmented-generationenterprise-ai

92% Correct. Without Handing It the Answers.

AgenticRAG enables AI models to autonomously navigate enterprise documents through iterative tool use, improving retrieval accuracy 5.9× over single-shot methods while reducing token costs.

Microsoft

May 16, 20269 minRead
agentic-aigovernance-frameworkrisk-management

Every AI Agent Audit Your Teams Run Is Missing the Same Thing

ARC Framework governs agentic AI through a capability lens that maps 46 risks to 88 technical controls using structured implementation guidance.

GovTech Singapore, Singapore University of Technology and Design

May 10, 20269 minRead
ai-coding-agentsrepository-context-filestoken-efficiency

The README That Makes AI Coding Agents 29% Faster

Adding AGENTS.md files to repositories reduces AI coding agent runtime by 29% and output tokens by 17% without changing task completion rates.

Singapore Management University, Heidelberg University, University of Bamberg, King's College London

April 24, 20269 minRead
voice-agentsred-teamingadversarial-evaluation

The Attack Your Voice Agent Cannot Be Trained Out Of

The first red-teaming framework designed not for AI models in isolation, but for voice agents as they are actually deployed

Fordham University, IBM Research

April 24, 20269 minRead
agent-safetyguardrailstest-time-adaptation

The Guardrail Blocking Half Your AI Agent's Legitimate Work

AGrail uses two collaborative LLMs with adaptive memory to defend AI agents against attacks, blocking 96% of threats while preserving 96% of legitimate actions.

The Ohio State University, University of Wisconsin-Madison, University of California, Davis

April 24, 202610 minRead
reinforcement-learningskill-libraryself-improving-agents

The AI Agent That Learns From Its Own Work

SAGE, a reinforcement learning framework from AWS and UW-Madison, trains AI agents to write and reuse programmatic skills across chains of related tasks, achieving 8.9% higher scenario completion than RL-without-skills baselines while using 59% fewer tokens.

University of Wisconsin–Madison; AWS Agentic AI (Amazon)

March 29, 202611 minRead
multi-agentretrieval-augmented-generationenterprise-search

77% Win Rate. This Is the RAG Architecture Behind It.

Atlassian's ADORE framework replaces single-pass RAG pipelines with a multi-agent, evidence-audited research loop that achieves a 77% win rate against ChatGPT Deep Research on business consulting tasks and outperforms all competitors on the DeepResearch Bench.

Atlassian

March 29, 20269 minRead
ai-governancerisk-managementconstitutional-ai

Your AI Vendor Has a Governance Problem

A Carnegie Mellon study reveals that Anthropic's Claude fails key transparency, bias, and accountability benchmarks under the NIST AI Risk Management Framework and EU AI Act, exposing significant governance gaps that enterprise buyers must audit before deployment.

Carnegie Mellon University, School of Computer Science, Privacy Engineering

March 17, 202610 minRead
agentic-aiprocess-automationprocess-mining

Your Workflow Automation Was Never Designed to Think

Researchers propose a five-layer architectural framework for Agentic BPM Systems (A-BPMS) that combines process mining, AI reasoning, and autonomous orchestration to move enterprise workflows beyond fixed rules into fully self-managing, self-optimizing operations.

University of Tartu

March 17, 20269 minRead
skill-orchestrationmulti-agentDAG-planning

Agent Skills Don't Compound. The Framework That Changes That.

AgentSkillOS shows that structured skill composition via tree-based retrieval and DAG orchestration dramatically outperforms flat skill provisioning, even when agents have access to identical tools.

Shanghai Artificial Intelligence Laboratory

March 17, 20268 minRead