Recent AI research delivers breakthroughs in coding, medicine, and autonomous systems. CodeMidas improves MiMo-V.5 on DeepSWE by 11.7%, while AutoRecLab achieves ~$1/run success. Medical prediction advances include MIST, enhancing survival prediction across four cohorts, and R-GEAN, scoring 0.464 on 240k admissions. Drug specificity boosts 84.8% via SpecOpt on 915 compounds, and clinical data accuracy improves by 27 and 20 pp using Ascent. Autonomous systems see RRDrive and PAANI advances, alongside a four-tier lab system aiding medical AI education.
Efficiency gains are substantial with L0-MoE offering 2.5x speedup, LazyAgent saving 42% CPU, and AgentRouter reducing costs by 72%. TinyCeNN-LM employs quality-gated attention, and AHRR achieves 2.6x chip design speedup. Designer-RSI increases design success from 72.7% to 99.3%, while Attention-Aware Routing improves GSM8K by 3.37 pp in MoEs. A population-supervised framework infers 3D red-cell properties with 0.86–0.98 correlation, and an affordable smart cane scores 0.82 F1 offline. Representation-guided learning yields 20-point classification and 13-point VQA gains.
Safety and alignment studies reveal critical gaps: TaxiGPT failures stem from map interference, and GameASG-Bench shows 55.3% task success versus 93.2% check pass. DUMA-Bench indicates attack success rising to 41.1%, while social influence reduces AI agent paper selection by 17.2% but raises subsequent rates by 45.55 pp. New benchmarks include ISA-Bench, FireWorldBench with 520 entries, and PhysAI-Bench with 10k instances. Limitations noted include benchmark saturation, task irrelevance, and annotation bottlenecks, with disagreements observed in underwater sonar and financial QA.
Methodological improvements involve tree-structured serialization fine-tuning, graph-guided navigation achieving 53.85% accuracy, and selective deferral for dementia crash prediction at 0.573 macro-F1. Social influence frameworks and consensus limitations show 0.21 error correlation. EvoPathBench demonstrates self-evolving agents often fail capability retention. A Nigerian fintech framework addresses governance gaps, and the Linear Representation Hypothesis is redefined as falsifiable. MAWILE audits LLM judges, and MSBD improves document segmentation. CaLR enhances diffusion reasoning, and TinyCeNN-LM uses quality-gated attention mechanisms effectively.
Key Takeaways
- CodeMidas improves MiMo-V.5 on DeepSWE by 11.7%.
- AutoRecLab achieves ~$1/run success in coding.
- MIST enhances survival prediction across four medical cohorts.
- R-GEAN scores 0.464 on 240k admissions.
- SpecOpt boosts drug specificity by 84.8% on 915 compounds.
- Ascent improves clinical data accuracy by 27 and 20 pp.
- L0-MoE delivers 2.5x speedup and LazyAgent saves 42% CPU.
- TaxiGPT failures stem from map interference issues.
- GameASG-Bench shows 55.3% task success vs 93.2% check pass.
- Designer-RSI increases design success from 72.7% to 99.3%.
Sources
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- AutoRecLab: Describe the Experiment, Get the Code!
- MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention
- Dual-Interest Sequential Product Recommendation With Multi-Granular SSM
- The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
- PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving
- GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
- World Modeling in Transformers
- GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development
- Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
- PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking
- Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks
- Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
- Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education
- LoRA Enhanced Contrastive Learning with SAS Vision Transformers
- PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation
- Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models
- Goal-driven Variant Categorization
- The Wisdom of Artificial Deliberative Crowds
- GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
- AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
- EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
- IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
- Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review
- Self-Organizing Agent Teams Learn to Reason Together
- Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA
- DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks
- Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems
- Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale
- Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model
- When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used
- ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures
- AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows
- A Compact Stance-Indexed Anterior-Posterior COP Representation for Parkinson's Disease Classification from Plantar VGRF
- R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction
- OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling
- LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs
- Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
- PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models
- From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving
- Tutoring Large Language Models to be Domain-adaptive, Precise and Safe
- Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
- WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
- On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks
- ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents
- Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control
- Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer
- LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning
- Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
- ACLArena: Agent Continue Learning in Multi-stage Post-training
- FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering
- Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
- Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search
- When More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search
- Incremental Consistency Execution for Autonomous Intelligent Systems
- Structured Decomposition for Reliable LLM-Generated Access Control Policies
- CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment
- APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
- Self-Healing Harness for Runtime Oversight of Agent Self-Modification
- A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories
- Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement
- How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation
- Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior
- VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
- LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning
- Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling
- Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
- AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction
- Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting
- Custom Named Entity Recognition and Topic Classification for Global Health Publications
- TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
- World State Generator
- DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
- Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents
- Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability
- Convex AI Compositionality and the Governance of AI System Populations
- Partner-Specific Affective Precision in Social Active Inference
- MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
- Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting
- SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models
- LIMIT: Less Is More for Instruction Tuning in Text-to-SQL
- EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
- DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
- UniK: Universal Knowledge Perception for Digital and Physical AI
- Divergent strategies and convergent outcomes in autonomous materials discovery
- Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems
- Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery
- Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models
- Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows
- Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events
- FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics
- Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees
- Harness-Zero: Harness Distillation via Agent-as-Harness
- Et Tu, Brute? Economic Misalignment in Personal AI Agents
- Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
- Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection
- GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes
- The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
- Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging
- Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
- Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information
- Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs
- When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
- PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
- TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents
- Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau
- CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
- CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy
- ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control
- Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework
- Generative Embodied Multiple Behavior Control Systems for Human-like Agents
- Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
- Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake
- Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models
- Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
- Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis
- Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation
- AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture
- CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
- A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning
- Calibrating Teacher--Student Discrepancy for On-Policy Distillation
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
- Offline Multimodal Large Language Models for Decision Support in Air Operations
- LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
- LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers
- Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving
- Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks
- Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection
- One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction
- GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation
- Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
- EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
- LLM-Generated Feature Pools for Time Series Anomaly Detection
- ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction
- A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
- Learning Cardiac Features: ECG Biometrics Across Time and~Exercise
- AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
- What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence
- SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity
- Can Agents Design Better Chips with a Higher Level Abstraction?
- TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers
- CaLR: Causal Latent Revision for Robust Diffusion Reasoning
- Attention-Aware Routing: Coupling Routing and Attention in MoEs
- Learning 3D biophysical cell properties from 2D images and cell-population statistics
- Social Influence and the Allocation of Scientific Attention in AI Populations
- An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users
- From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis
- Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
- MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
- Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
- Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
- Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
- Representation-guided in-context learning for medical image interpretation with multimodal large language models
- Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech
- A Survey on the Linear Representation Hypothesis
- Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Comments
Please log in to post a comment.