Paper Review

Paper: ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Listen to this article.

Only the latest audio is kept; older files are removed on each update.

Problem

Text-to-image (T2I) models are fantastic at generating images, but they struggle with complex tasks requiring real-world knowledge and multi-step reasoning. Current approaches to giving these models “agent” abilities—allowing them to act more intelligently—either have rigid workflows or only control parts of the image generation process. This means the various steps (reasoning, using external tools, and generating images) aren’t working together as effectively as they could be.

Paper: ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Training search agents that need to perform complex, multi-step tasks – like retrieving information and reasoning over it to answer questions – is tricky. Existing methods often treat every action the agent takes during a search equally, whether it leads closer to the right answer or not. This means valuable actions can get lost in the noise of less helpful steps, hindering learning.

Paper: HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Training robots to perform manipulation tasks (like picking up and moving objects) often struggles with a lack of good data. Collecting accurate, high-quality data directly from real robots can be expensive and time-consuming. While data collected without a robot (“UMI” data - Unimaged Manipulation) is easier to scale, it’s typically used only for initial training and then fine-tuned on a small amount of real robot data. This paper challenges that approach by asking: what if we could make UMI data so good that we didn’t need the expensive real-robot portion at all?

Paper: JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Creating truly useful AI-powered creative tools requires more than just generating assets on demand. Current systems like prompt-based or chat-based generators often treat each request in isolation, failing to maintain context, track revisions, or manage the complex workflow of a real-world creative project (e.g., video editing, graphic design). Commercial “creative agent” systems exist but are largely closed off, hindering research into how they actually work and make decisions.

Paper: Kimi K3: Open Frontier Intelligence

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Training increasingly large language models (LLMs) has become computationally expensive and inefficient, hindering progress in the field. Existing architectures struggle to effectively utilize all parameters during inference, and scaling these models can lead to diminishing returns. This paper tackles that challenge.

Method

The authors introduce Kimi K3, a 2.8 trillion parameter Mixture-of-Experts (MoE) model aiming for more efficient scaling. Key components of their approach include:

Paper: AREX: Towards a Recursively Self-Improving Agent for Deep Research

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Deep research is challenging because finding potential solutions often takes significant effort, while checking whether those solutions meet all the required constraints (multiple criteria) can be broken down into smaller, more manageable steps. This “discovery-verification asymmetry” creates a bottleneck: simply searching for longer doesn’t necessarily lead to better results.

Method

The paper introduces AREX, a family of “Recursively Self-Improving” (RSI) deep research agents designed to address this challenge. AREX operates with an alternating two-loop structure:

Paper: SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Training massive, trillion-parameter Mixture of Experts (MoE) language models like DeepSeek-V4 presents significant engineering challenges when using distributed training systems. The paper highlights issues including intense memory usage, communication bottlenecks, and inefficient processing during the post-training phase—specifically, Full Parameter Post-Training (CPT) and Supervised Fine Tuning (SFT). While most existing solutions rely on GPU clusters, this research explores an alternative approach leveraging Ascend Neural Processing Units (NPUs).

Paper: VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Current open-source video understanding models face several limitations. They often struggle to generalize across different types of videos, performing well only in specific niches. These models also tend to be computationally expensive and may not be fully accessible for researchers or developers, with key training details and datasets withheld.

Method

The paper introduces VideoChat3, a “fully open” video-centric Multimodal Large Language Model (MLLM) designed to overcome these limitations. The core approach combines two key elements:

Paper: Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Den...

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Current benchmarks used to evaluate AI agents often focus on simple tasks that complete quickly and are judged solely by their final outcome. This doesn’t give a full picture of an agent’s capabilities, especially when dealing with complex, real-world scenarios requiring sustained effort and iterative problem-solving. Existing “terminal” benchmarks (which judge only the end result) provide limited insight into intermediate progress and partial solutions due to sparse reward signals.

Paper: UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Evaluating proactive AI agents—those designed to operate tools and assist users in real-world environments like personal assistants or automated workflows—is currently difficult. Existing benchmarks often use simplified, sandboxed testing grounds and evaluate agents only on single interactions. Additionally, these benchmarks categorize tasks in ways that blur the lines between different underlying capabilities of the models, making it hard to pinpoint why an agent succeeds or fails.