Research Brief
Key Takeaways
- • Key findings from research papers
Sources
- GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
- TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
- When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
- H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System
- When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs
- Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching
- OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents
- Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation
- Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
- Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search
- CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
- CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
- PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data
- Emotion in an active inference model of human driving
- Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains
- Towards an Argumentative Foundation for Evaluative AI
- The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
- Towards Researcher Agents for Knowledge-Graph Question Answering
- TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
- When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
- MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
- Controlled Memory Interference in Continual LLM Agents
- TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair
- The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
- An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop
- Protecting patient privacy in clinical foundation models: Technical and legal perspectives
- Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns
- IntelliAudit: Using Large Language Models to Evaluate Audit Controls
- An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
- AndroidReality: How Far Are Mobile Agents from the Real World?
- Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
- QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing
- Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
- SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
- GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
- Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
- ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
- Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs
- SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
- Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning
- KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs
- Legal Responsibilities Using Autonomous Agents For Artificial Intelligence
- Thought-Level Beam Search for Reasoning
- VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge
- Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
- SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution
- Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE
- SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents
- Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
- Generative Models: Principles, Architectures, and Applications
- Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework
- Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation
- TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance
- Constraining ontology mappings using metaphysical choices
- Quantization Degradation in Large Language Models: A Signal-Noise Perspective
- Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges
- Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy
- Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue
- A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
- Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios
- Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?
- StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning
- Mitigating Over-Personalization in LLMs via Structured Memory
- CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
- SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance
- TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation
- MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
- HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails
- Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production
- Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation
- Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
- Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
- What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files
- Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing
- Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
- SDDBMs: Soft Denoising Diffusion Bridge Models
- FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents
- VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
- Deep probabilistic logic programming for diagnostic reasoning from incomplete information: A case study in stroke detection
- ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
- Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
- MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
- Can Open-Weight Models Compete on Financial Text Comprehension?
- A QUBO-Inspired Computational Framework for Airport Landside Bottleneck Diagnosis and Dynamic Dispatch Optimization
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration
- The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task
- Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
- Matryoshka Language Model Suites
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
- PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
- Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models
- Three Generations of Healthcare IT: From the Digital Record to the Computable Care Process
- Improving Generalization Robustness of Multimodal RLVR
- LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
- Full-bandwidth transformer
- AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
- AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups
- Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration
- Depth-Aware Implicit Neural Representation Priors for 3D Gravity Inversion
- Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis
- Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging
- Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
- DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning
- PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs
- ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models
- RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
- MELLON - Multimodal Enhanced LLM for Online Navigation
- Motif 3: Technical Report
- RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
- When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
- CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
- SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
- FemWear: A Specialized Wearable Foundation Model for Women's Health
- Agentic Router: An Execution-Grounded Continual Learning Approach With Memory
- From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
- CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
- CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation
- Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models
- CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
- Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
- Entropy-based Code Adversarial Translation for Real-world Repository Migration
- MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
- Linearized 2-Simplicial Attention
- ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
- OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks
- CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits
- Control-Oriented Scenario Tree Construction through Reinforcement Learning
- GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models
- CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
- Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
- From Prompt to Harness: Coderlet from Scratch
- Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents
- verdi: retrieval is not transfer for continual world model optimization
- One Adapter Pair per Model: A Universal Activation Interface for Language Models
- Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
- CoRCi: Cross-Reconstruction of Coherent Interests Modeling in Cross-Domain Sequential Recommendation
- From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization
- Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
- Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
- Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
- DSLE: A Learning Environment for Dark Souls Boss Encounters
- ArchAgent v2: A Case Study with the Data Prefetching Championship
- Towards Expert-level Medical AI for Real-time Video Consultations
- Agentic Auto-Research is Fuzz Testing
- Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity
- KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
- LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling
- ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management
- An Explainable GNN Framework for Component-Level Anomaly Diagnosis
- SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge
- Signature-Guided Capacity Occupancy for Dense Expert Merging
- Structure-Preserving Uncertainty Propagation in First-Order Proof Search
- TRACE: TRajectory Attribution for Automated Context Engineering
- Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
- CoRe-UIE: Rethinking Coexisting and Region-wise Degradation for Underwater Image Enhancement
- Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
- Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection
- From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings
- FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
- AI Evaluation Should Measure Verification Cost, Not Correctness Alone
- SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
- CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
- Mismatch Matters: On-Policy Distillation Beyond Token Agreement
- ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs
- ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
- The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games
- Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents
- Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
- A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition
- PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
- EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
- Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
- SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests
- Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata
- Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization
- Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs
- TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
- LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
- Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
- Estimating Uncertainty in Galaxy Morphology Classification
- CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
- Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
- LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving
- Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning
- Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders
- A Minimal $\kappa$--$\tau$ Logic for Risk-Sensitive Abduction
- Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets
- Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
- A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning
- JustLLMGRPO: Radiographic Control for Chest X-Ray Generation
- The Authority Expectancy Effect in Multi-User Conflict
- CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity
- TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
- GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
- Back to the Future: A workbook time machine for spread sheet creation benchmarks
- Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge
- CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
- Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
- Contextual Value Alignment via Multilayer Combinatorial Fusion
- From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems
- Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature
- Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation
- P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation
- Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
- Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
- Walking through Discussions: A Mobile Visual Analytics System for In-Situ Group Discussion Analysis
- Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
- Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
- UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
- LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems
- Metanormative Theory for RL-Based Moral Agents
- Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment
- Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models
- Improving Constraint Models with LLM Agents
- Neurosymbolic Discovery of Algebraic Graph Constructions
- REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
- The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
- Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
- Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems
- NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
- Training Variable Long Sequences with Data-Centric Parallel
- Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility
Comments
Please log in to post a comment.