Files
val-blog/org/cases/arxiv_digest/output/latest.md
T

63 lines
5.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ArXiv Daily Brief - 2026-03-14
## 🧠 今日 Top 3(中文可读版)
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- arXiv: http://arxiv.org/abs/2603.12266v1
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- arXiv: http://arxiv.org/abs/2603.12262v1
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- arXiv: http://arxiv.org/abs/2603.12238v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- arXiv: http://arxiv.org/abs/2603.12266v1
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...cs.CV
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- arXiv: http://arxiv.org/abs/2603.12262v1
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneouslycs.CV
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- arXiv: http://arxiv.org/abs/2603.12238v1
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generationcs.CV
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
- arXiv: http://arxiv.org/abs/2603.12249v1
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoningcs.CL
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
- arXiv: http://arxiv.org/abs/2603.12246v1
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Trainingcs.AI
## 🆕 最新上新 Top 10
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。