Recent advances in LLM efficiency and reliability include Harness Compilation, which improves VLM scores by 9.9-23.9 points, and LLoCoT, a looped latent-reasoning framework reducing latency by ~36x. Plan-and-Patch demonstrates diffusion models outperform autoregressive ones in plan repair (53.7% vs 27.0%) while cutting latency by 39-46%. RaReCache enables cross-model KV cache reuse, retaining 95-99% accuracy across large parameter gaps, and Memento 3 allows frozen agents to learn world models, clearing all 25 ARC-AGI-3 games with 100% RHAE. OnTrack detects failures in ~1ms to save 18% compute, and EvoAlloc reduces evaluation costs by 59-82% via self-evolving resource allocation.
Agent safety and reliability remain critical, with typed decision models showing unreliable accuracy (36%-72%) against defenses like prompt injection. Legal reasoning models exhibit disconnects between citations and verdict dependence, while HarmBench fails to measure single 'harmful refusal' attributes due to saturation. Safety studies reveal 56.92% of trajectories contain unfulfilled obligations, prompting ObligationGuard with 56.52% recall. Ecological safety theory suggests misaligned populations face critical takeoff thresholds, and BrickBench finds agents satisfy physical constraints but fall short of human design quality in LEGO tasks.
Methodological innovations address calibration and reasoning gaps. Bayesian neural networks achieve efficient inference by dynamically allocating Monte Carlo samples, reducing latency while maintaining error guarantees. TRACE traces emotions in social scenes to reveal model gaps, and ReCast improves failure attribution by learning step representations for best Hit@1 performance. Universal Textual Teaching raises student accuracy on math and code tasks via parameter-free textual primers, and GeoReform boosts geometry reasoning accuracy from 42% to 56% by treating formalization as an optimizable policy.
Domain-specific applications show significant progress. Synthetic vascular audits predict equivalence within 4.0e-15, though noisy anchors constrain calibration. OA-MAP achieves 0.80 AUROC for knee osteoarthritis structural progression using multimodal data. AtomWorld-Mirror accelerates materials simulation 1000-10000x via macro-step inference, and BEVPIPE enables portable BEV perception deployment with 19.5x speedup. TokenBank uses structured contracts to manage AI service costs, reducing mean expenditure by $304.88, while Agent-controlled forgetting cuts token usage by 50% and costs by ~67% in tool-use scenarios.
Key Takeaways
- Harness Compilation improves VLM scores by 9.9-23.9 points.
- LLoCoT reduces latency by ~36x via looped latent reasoning.
- Diffusion models outperform autoregressive ones in plan repair (53.7% vs 27.0%).
- RaReCache retains 95-99% accuracy across large parameter gaps.
- Memento 3 clears all 25 ARC-AGI-3 games with 100% RHAE.
- Typed decision models show unreliable accuracy (36%-72%) against defenses.
- HarmBench fails to measure single 'harmful refusal' attributes due to saturation.
- GeoReform boosts geometry reasoning accuracy from 42% to 56%.
- AtomWorld-Mirror accelerates materials simulation 1000-10000x.
- Agent-controlled forgetting cuts token usage by 50% and costs by ~67%.
Sources
- Complexity of Grounded Semantics and Preferred Semantics in Finitary Argumentation Frameworks
- An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation
- Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?
- Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
- TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
- Why machines will still not rule the world
- From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
- SpikeSSL: A Universal Spike Inference Framework with Dynamics-Informed State-Space Layers
- Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
- An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment
- The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate
- Plan-and-Patch: Diffusion Language Models for Agentic Planning
- Reading the Room: Foundations, Design, and Challenges of Normative Competence in LLMs
- Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
- Distillation for Incrimination and Distillation for Capabilities
- When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
- What to Admit and How to Present: Governing Persistent Memory in LLM Agents
- Social Pain Disrupts Emotion-Action Brain-State Dynamics in Adolescents with Non-Suicidal Self-Injury
- When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
- Open-ended Scientific Discovery with Possibilistic Reasoning
- DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
- ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems
- RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
- RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
- UniData: Universal Multimodal Instruction Generation Pipeline
- Estimating great expectations under autoregressive language models with potentials
- Cognition-Oriented Emotion Tracing from Causes to Consequences in Real-World Social Scenes
- Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
- Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
- BridgeGuard: Explicit Safety Drift for Diffusion-based Autonomous Driving
- Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds
- Workerville: Towards an Organizational Behavior Account of Agent Safety
- AgentEvolver: System-Wide Self-Evolution Through Task Execution
- Evidence-Traceable Dynamic Interviewer Architecture for Expertise-Adaptive Qualitative Interviews Using Local LLMs
- Intervention anchors and scientific verification in synthetic vascular predictive representations
- DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents
- Scalable AI Uncertainty Quantification via Generalized Laplace Active Subspaces
- Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models
- What Output-Only Review Cannot Verify: Study Contracts for Research Agents
- MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
- MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation
- InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers
- Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models
- An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under Resampling
- Learning Probabilistic Logic Programs with Functional Gradient Guided Language Models
- One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
- When Has a Bayesian Neural Network Sampled Enough? Adaptive Inference Time with Statistical Guarantees
- Recursive Self-Improvement through Multi-Agent Self-Supervision
- Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
- Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution
- Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
- HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
- Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
- On the estimation and validity of AI time horizons---a statistical look at the METR plot
- BrickBench: Evaluating Agentic Brick Design
- Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
- GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
- OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
- HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
- Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
- Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models
- Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization
- Q-Shaped Options for Hierarchical Reinforcement Learning
- OA-MAP: Evidence-Grounded Multi-Agent Multimodal Framework for Interpretable Knee Osteoarthritis Progression
- Use and Disuse: Intent-Structured Experience Consolidation for Memory and Learning in LLM Agents
- Universal Textual Teaching for LLMs
- EvoAlloc: A Self-Evolving Resource Allocation Agent for Efficient Program Evolution
- Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes
- Probability-Signature Dynamics: Unpacking Modular Addition Learning Within Two-Layer Networks
- Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
- RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
- Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding
- Onboard Marine Anomaly Detection on $\Phi$sat-2: From Simulation-Based Development to In-Orbit Demonstration
- MemTrial: Learning When to Trust Memory in LLM Portfolio Agents
- MultiWorldBench: Do Independently Controlled Views Describe One Shared World?
- Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent Systems
- Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies
- Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
- ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
- AtomWorld-Mirror: Macro-Step World Modeling of Critical Evolution Backbones for Materials Dynamics
- The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute
- RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees
- Writing for the Reviewer: Defensive Writing in GPT Models
- SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
- EvoSim: Learning to Model, Modeling to Learn
- TokenBank: Financial Infrastructure for AI Services
- Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild
- AliO: Output Alignment Matters in Long-Term Time Series Forecasing
- GameCommBench: A Unified Benchmark and Type-Aware Evaluation for AI-Generated Game Commentary
- Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL
- OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
- AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
- Agent-Controlled Forgetting for Tool-Using Agents: Reversible Context Curation in Practice
- Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction
- Speaking the Navigator's Language: Trajectory-Grounded Instruction Translation for Frozen Aerial VLN Agents
- Verification and Self-Improvement in Agentic AI: Foundations and Limits
- Prior or Feedback? What an LLM Uses When Adapting Neural Operators
- Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
- A Structural Theory of Cognitive Representation and Problem Solving,Contexts, Invariance, and the Knowledge Space
- Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance
- How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices
- LEVER: Adaptive Cost-Aware Proof Search Over AND/OR Graphs
- Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
- A 3D Characterization Framework for Intelligent Sequential Decision Making
- Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
- Lamarck's Driving School: Discovering Autonomous Driving Training Strategies through Evolutionary Competition
- Harness Evolution Hits a Ceiling: When Weight Training Should Begin
- Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
- Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models
- Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference
- MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction
- DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
- ORDO: Operation-level Round-aware Dynamic Ordering for MIP Presolve
- TaReD: Tool-Aware Recursive Decomposition for Long-Horizon Tasks
- LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models
- Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored Feedback
- How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
- Learning How to Search for Plans with Exponentially Less Space
- StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
- On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
- Whose Ground Truth? Embracing Ambiguity in Human-Centered AI
- When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation
- Structure Tax: How Structured Output affects LLMs Performance
- PulseBound: Future-Beat State Forecasting Under an Explicit Information Boundary
Comments
Please log in to post a comment.