Recent advances in agentic AI demonstrate significant gains in trading, safety, and efficiency. EvolveTrade improves LLM trading agent returns via policy refinement, while OBC-Prune enhances LLM reasoning pruning by calibrating causal importance of tokens. In safety and verification, Publication Authority introduces a PAC-2026 protocol for falsifiable AI records, and GuardEn boosts visual safety assessment by 9.8 F1 points using executable rule entailment. Autonomous driving policies trained via imitation learning enable collision-free operation, and RiskWorld achieves the lowest collision rates using flow-guided occupancy evolution.
Critical gaps in reliability and evaluation protocols remain a primary focus. ERPBench reveals general GUI agents succeed in only 3% of enterprise record-saving runs, whereas AutoTuneBench exposes how naive baselines can manufacture speedups. Frontier models show distinct evidence-acquisition policies driven by severity, yet Chain-of-thought monitoring fails to detect algorithmic collusion in pricing agents. To address these issues, TuiML offers a self-contained ML library with native validation, and an independence-graded audit protocol argues binary checks fail to detect collusion across principal, substrate, and evidence axes.
Efficiency and specialized applications drive further innovation. Inference optimizations identify a Pareto frontier where FP8 weights retain 99.4% accuracy with reduced latency, and Edge0 serves 35B MoEs at 20tok/s within 3GiB memory. Graph-based RAG reduces costs by 57% while maintaining quality, and EffiRAG demonstrates similar cost reductions. In specialized domains, HPOQuest improves rare-disease diagnosis recall by up to 30%, and a Finnish Turing Test shows LLMs pass when cultural context is accounted for. Finally, collective loss of control in agent systems is explained as an epidemic of mutation where local deviations propagate into systemic failure.
NLP and biomedical applications highlight both progress and limitations. Mahalanobis-Ensemble Decoding enhances semantic diversity with negligible overhead, and MAGER enables frozen LLMs to detect fake news via meta-path discovery. However, SNOMED CT concept recommendation shows sparse TF-IDF outperforms other methods but drops significantly for rare concepts. Knowledge editing suppresses rather than erases original facts, and infinite-parameter LLMs generate weights from live data via Bayesian hypernetworks. Education studies link early GenAI reliance to negative impacts, while governed systems outperform hosted services in normative document QA.
Key Takeaways
- EvolveTrade improves LLM trading agent Sharpe Ratio and cumulative returns.
- Publication Authority introduces PAC-2026 protocol for falsifiable AI records.
- Evidence masking boosts compositional generalization accuracy by median 0.846–0.859.
- Physics-constrained digital twins detect stealthy false data injection in urban flow.
- SSLD matches baselines on 1600+ graph coloring instances despite 195x runtime cost.
- COTQ evaluation shows structural consistency with ESA WorldCover for land cover.
- GVD unifies versioning and deduplication reaching 0.97 F1 without LLMs.
- FairCompressAgent reduces FPGA storage by 59.54% while improving precision.
- ERP Bench reveals GUI agents save records correctly in only 3% of runs.
- FP8 weights retain 99.4% accuracy with reduced latency in inference optimizations.
Sources
- EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
- Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records
- What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization
- Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees
- One Color Preprocessing Improves DSATUR
- A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products
- GVD: Governed Versioning and Deduplication for Document Repositories
- Imitation Learning for Autonomous Driving in CARLA
- FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment
- A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning
- Learning Heterogeneous Preferences
- OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
- ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
- Do Frontier Models Seek Safety Evidence Before Acting?
- The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?
- TuiML: Machine Learning for AI Agents
- Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits
- When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI
- Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI
- Teaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools
- The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
- Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning
- Symbolic Temporal Supervision of LLM Agents Using Contracts
- Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost
- AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
- When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation
- Building Trust in Artificial Intelligence: A Necessity for Railway Applications
- Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI
- BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
- REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement
- Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
- Visual Compliance via Executable Safety Rule Entailment
- What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models
- TRIPROBE: Probing Task Separability Beyond Classification for XAI
- HPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition
- Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland
- Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
- Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery
- The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models
- Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving
- The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses
- Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making
- CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
- Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale
- Clueing up LLMs with Tool-Augmented Deductive Reasoning
- Which LLM is Best for Translating Natural Language Goals to PDDL
- Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization
- Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation
- Suppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing
- Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
- Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
- MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
- Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation
- Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum
- Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models
- AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution
- RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
- Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations
- Collaborative Memory for Multi-Agent VLM Systems
- GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents
- CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
- Flag Game: A Toy Model for Mechanistic Swarm Interpretability
- Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning
- Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
- Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs
- First Token Matters: Understanding Safety Collapse in Large Reasoning Models
- WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories
- Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment
- WFM: Wiki Foundation Model for Complex Agentic Reasoning
- Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting
- Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition
- NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation
- Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents
- Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
- Missing Bridges: Composition-Aware Active Imitation Learning
- SNOMED CT Concept Recommendation from Masked Clinical Context
- Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
- Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection
- SAGE: Governed Artifact Generation from Enterprise Guidelines
Comments
Please log in to post a comment.