Researchers have developed a new architecture for reliable LLM-Powered Decision Engines in large-scale supply chain operations. The architecture combines large language models with mathematical optimization, probabilistic forecasting, and safety-constrained decision filtering. It has been tested on real-world data and shown to improve performance and safety guarantees compared to traditional rule-based and optimization-only systems.
A new benchmark, SIMGUIDE, has been introduced for evaluating personalized AI agents. The benchmark assesses an agent's ability to treat users as single entities and adapt to different priorities across life contexts. It has been tested on a dataset of 47 preference-conditioned planning tasks and shown to outperform existing baselines.
A new framework, CaSKG, has been proposed for counterfactual-causal skill graph construction. It uses a hierarchical multimodal MoE for interstitial lung disease classification and has been shown to outperform existing methods on a dataset of 138 unique artifacts.
A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees. It has been tested on a dataset of 100k training windows and shown to outperform existing methods in terms of accuracy and efficiency.
A new framework, SIMGUIDE, has been introduced for evaluating personalized AI agents. The benchmark assesses an agent's ability to treat users as single entities and adapt to different priorities across life contexts. It has been tested on a dataset of 47 preference-conditioned planning tasks and shown to outperform existing baselines.
A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees. It has been tested on a dataset of 100k training windows and shown to outperform existing methods in terms of accuracy and efficiency.
A new framework, SIMGUIDE, has been introduced for evaluating personalized AI agents. The benchmark assesses an agent's ability to treat users as single entities and adapt to different priorities across life contexts. It has been tested on a dataset of 47 preference-conditioned planning tasks and shown to outperform existing baselines.
A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees. It has been tested on a dataset of 100k training windows and shown to outperform existing methods in terms of accuracy and efficiency.
A new framework, SIMGUIDE, has been introduced for evaluating personalized AI agents. The benchmark assesses an agent's ability to treat users as single entities and adapt to different priorities across life contexts. It has been tested on a dataset of 47 preference-conditioned planning tasks and shown to outperform existing baselines.
A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees. It has been tested on a dataset of 100k training windows and shown to outperform existing methods in terms of accuracy and efficiency.
Researchers have developed a new architecture for reliable LLM-Powered Decision Engines in large-scale supply chain operations. The architecture combines large language models with mathematical optimization, probabilistic forecasting, and safety-constrained decision filtering. It has been tested on real-world data and shown to improve performance and safety guarantees compared to traditional rule-based and optimization-only systems.
Key Takeaways
- Researchers have developed a new architecture for reliable LLM-Powered Decision Engines in large-scale supply chain operations.
- A new benchmark, SIMGUIDE, has been introduced for evaluating personalized AI agents.
- A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees.
- A new framework, SIMGUIDE, has been introduced for evaluating personalized AI agents.
- A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees.
- Researchers have developed a new architecture for reliable LLM-Powered Decision Engines in large-scale supply chain operations.
- A new benchmark, SIMGUIDE, has been introduced for evaluating personalized AI agents.
- A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees.
- A new framework, SIMGUIDE, has been introduced for evaluating personalized AI agents.
- A new method, PhaseShift, has been introduced for allocating measurement budgets in quantum learning with finite-shot generalization guarantees.
- Researchers have developed a new architecture for reliable LLM-Powered Decision Engines in large-scale supply chain operations.
Sources
- Reliable LLM-Powered Decision Engines for Large-Scale Supply Chain Operations: Architecture, Safety, and Performance Guarantees
- SIMGUIDE: Procedurally Grounded Multi-Context Representations for Personalized Agent Planning
- Measurement-Budget Allocation in Quantum Learning with Finite-Shot Generalization Guarantees
- Solving Robust POMDPs with Omega-regular Objectives via Partially Observable Stochastic Games
- FrontierChallenge: Evaluating Scientific Workflow Completion
- post-graph-rag: A PostgreSQL-Native Graph RAG Engine
- Semantic Graph Unification for Industrial Digital Threads: Bridging 11 Heterogeneous Manufacturing Systems Through Ontology-Driven Knowledge Graphs
- Natural Language Input, Semantic Track Representation, and LLM Inference: Making the Maritime Information Exchange Model Tractable
- CVE-SAI: Counterfactual Visual Evidence-Guided Selective Attribute Indexing for Risk-Controlled E-commerce Search
- Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs
- Auto-Policy, not Auto-Skill: Compiled Agent Skills for the Physical World
- LifePlanner: Evaluating LLM Agents for Geo-spatial Planning with Social Media Data
- Hierarchical MoE for Multi-Modal ILD Diagnosis
- PhaseShift: Topology-Aware Data Harmonization and Model Consolidation Across Signalized Intersections
- BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks
- LLM-Driven, Datasheet-Aware Automated Hardware Compatibility Verification for Early-Stage, Pre-Schematic Embedded System Design
- FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review
- Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory
- Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs
- Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
- CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
- PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning
- LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents
- Narcissus: Program Synthesis Using Context-Aware LLM Approximations
- Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
- How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation
- ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs
- AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
- Imitation Learning for Connection-Tableau Construction
- SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
- Quantitative Analysis of $\omega$-Regular Robust MDPs
- Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems
- Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
- SwarmWorld: Stigmergic technological evolution in societies of language-model agents
- Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures
- Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions
- ToST: A Tree-of-Thought Socratic Teaching Framework for Multi-Path Guidance and Parallel Thinking
- Using profiles of cognitive capability to assess AI suitability for workplace tasks
- FLARE: Verifying MILP Reformulations with LLM-Based Theorem Proving
- Tunable Tool-Call Rates in LLM Agents via Representation Steering
- FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
- PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
- SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
- AI-Powered Mental Health Chatbots in Africa: A Systematic Review and Culturally Adaptive Framework
- VLM-based automatic multi-granularity graph representation of building layouts for design informatics
- Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
- LivingRAG: Augmenting Graph RAG with Experience
- Candidate supply and answer selection shape the value of LLM judging in multi-agent systems
- Training Alignment Auditors via Reinforcement Learning
- Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness
- Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
- Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks
- Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2
- RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
- Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning
- Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows
- Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
- Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
- In-Context Inpainting for Time Series Forecasting
- Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems
- Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment
- PROOF-Gen: From Optimized Data to Better Distillation
- BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification
- MARS: Multi-Specialist LLM Relay System for Competitive Programming
- Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining
- Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors
- More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
- Recursive Agentic Reasoning
- More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving
- Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction
- Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning
- Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling
- When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
- Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design
- Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
- Reflection with Action-Induced Visual Differences for Desktop GUI Agents
- PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents
- Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems
- Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models
- ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
- EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals
- AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
- Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
- Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
- Task-Adaptive Rubrics for GUI Reward Modeling
- OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
- AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL
- Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
- STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation
- TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
- Evaluating Multiple LLM Generations with Validated Task Coverage
- Constraint-Guided Enterprise Data Mapping with Large Language Models
- RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
- Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding
- Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
- Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
- VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models
- Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs
- Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
- Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning
- Lifted Model Construction under Approximate Commutativity
- A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
- ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
- Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling
- Mahalanobis-Based Multi-Head Attention for Complex State Propagation
- EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents
- Discovering Adaptive Transmission Programs for Collective Innovation
- PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
- Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms
- Causal Modelling of Support Interventions for Student Competency Assessment
- Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
- The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
- Confident at the moment of action: belief miscalibration in LLM play under hidden information
- Meta$^n$: Recursive Self-Improvement through Emergent Depth
- Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
- Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA
- Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core
- StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
- CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
- Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
- FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
- A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads
- Partial Identification under Causal Orders by Linear Programming
- From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
- The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
- SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception
- SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction
- MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG
- Preference Data Selection for Mitigating the Alignment Tax in Large Language Models
- Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
- Relative Time Intervals Representation for Word-level Timestamping with Masked Training
- Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing
- Provenance Guided Incremental Learning Under Evolving Concept Definitions
- AI Finds A Way
- Exploit More, Explore Smarter for Budget-Constrained Agentic Search
- SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
- Joint Optimization of Tool Creation and Use for Large Language Model Agents
- When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
- Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites
- HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning
- Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
- AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
- Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
- OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning
- ReproAgent: Contract-Guided Paper-to-Code Reproduction
- Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling
- Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
- How much of a measured AI preference is the model, and how much is the instrument?
- Function-Level Execution Feedback for Code Preference Optimization
- LLM Agents Perform Controlled Experiments Using Simulation Models
- RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
- Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks
- ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
- Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
- TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery
- A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
- MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models
- FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare
- AI Agents Push Humans Out of the Loop
- Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
- Do LLMs Understand Limit Order Book Dynamics?
- Automata from Agent Traces: Failure and Next-Step Prediction
- Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
- A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
- Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search
Comments
Please log in to post a comment.