CATArena Advances AI Agent Testing While Denario Simplifies Financial Research

Recent research demonstrates that targeted supervised fine-tuning on persuasive data increases LLM persuasiveness, while preference optimization offers no additional benefit. Agentic capabilities advance significantly through methods like RECAST, which routes evidence via computation and retrieval to achieve 75.6% success on six benchmarks, and TRACE, a constraint-tree algorithm reaching 86% success on RecMovie. Coding-agent performance can be predicted via base-model screens matching post-trained SWE-bench results, and heterogeneous GNNs reduce multi-agent path planning runtime by 40%.

Safety and reliability remain critical challenges, with theoretical bounds proving perfect agents cannot guarantee safety in partially observed environments. On-device safety is fragile, where modifying 0.19% of weights yields 53% Basic ASR, and standard task success masks safety failures, increasing collisions by 12.3x in real-time execution. Adaptive reasoning budgets balance latency and transparency in guardrails, while separating co-witnesses reduces unsafe task completion by 33.59 percentage points. Refusal circuits identified by SafeEvo reduce harmfulness by 63.21% and over-refusal by 58.44%.

Long-horizon learning benefits from self-evolving reward adaptation and dynamic computation optimization via adaptive KV caching, reducing FLOPs by up to 30%. Knowledge accumulation frameworks like Knowledge Weaver and UniSkill improve agent performance by curating distinct, reusable skills without redundant overlap. Humanize employs mechanical judgement engineering for high success in agentic coding, though stopping remains a weakness, while CADFather autonomously reconstructs CAD models by coordinating tools without additional training.

Evaluation benchmarks reveal gaps in end-to-end urban diagnosis workflows and tool-using agents often ignore silent failures at 58.8% rates compared to 91.3% for explicit errors. Clinical suicide risk assessment improves via ontology-grounded GraphRAG, outperforming vector RAG in completeness and relevance. Retrieval-Augmented Models show state-change magnitude predicts source reliance better than latent trajectory shifts, and private holdouts generally narrow score-satisfaction gaps but widen them in specific cases.

Key Takeaways

  • Targeted supervised fine-tuning increases LLM persuasiveness; preference optimization adds no benefit.
  • RECAST achieves 75.6% success on six benchmarks by routing evidence via computation and retrieval.
  • TRACE constraint-tree algorithm reaches 86% success on RecMovie using language feedback.
  • Theoretical bounds prove perfect agents cannot guarantee safety in partially observed environments.
  • On-device safety is fragile; minor weight modifications yield 53% Basic ASR in LLaMA-2-7B-Chat.
  • Standard task success masks safety failures, increasing real-time execution collisions by 12.3x.
  • Adaptive reasoning budgets balance latency and transparency in safety guardrails effectively.
  • Separating co-witnesses reduces unsafe task completion by 33.59 percentage points.
  • SafeEvo identifies refusal circuits, reducing harmfulness by 63.21% and over-refusal by 58.44%.
  • Tool-using agents ignore silent failures at 58.8% rates versus 91.3% for explicit errors.

Sources

NOTE:

This news brief was generated using AI technology (including, but not limited to, Google Gemini API, Llama, Grok, and Mistral) from aggregated news articles, with minimal to no human editing/review. It is provided for informational purposes only and may contain inaccuracies or biases. This is not financial, investment, or professional advice. If you have any questions or concerns, please verify all information with the linked original articles in the Sources section below.

ai-research machine-learning arxiv research-paper llm-persuasiveness recast trace safety-in-artificial-intelligence on-device-safety adaptive-reasoning-budgets safe-evo

Comments

Loading...