# ArXiv Daily Brief - 2026-03-14 ## 🧠 今日 Top 3(中文可读版) 1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning** - 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning - 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep... - 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning. - 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de... - 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de... - arXiv: http://arxiv.org/abs/2603.12266v1 2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously** - 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously - 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction. - 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding. - 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la... - 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction. - arXiv: http://arxiv.org/abs/2603.12262v1 3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation** - 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成 - 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation. - 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation. - 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi... - 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation. - arXiv: http://arxiv.org/abs/2603.12238v1 ## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索) 1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning** - arXiv: http://arxiv.org/abs/2603.12266v1 - 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue - 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...(cs.CV) 2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously** - arXiv: http://arxiv.org/abs/2603.12262v1 - 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang - 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously(cs.CV) 3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation** - arXiv: http://arxiv.org/abs/2603.12238v1 - 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu - 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation(cs.CV) 4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning** - arXiv: http://arxiv.org/abs/2603.12249v1 - 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang - 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning(cs.CL) 5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training** - arXiv: http://arxiv.org/abs/2603.12246v1 - 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su - 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training(cs.AI) ## 🆕 最新上新 Top 10 1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1 2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1 3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1 4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1 5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1 6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1 7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1 8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1 9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1 10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1 ## Val 今日建议 - 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。 - 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。