63 lines
5.7 KiB
Markdown
63 lines
5.7 KiB
Markdown
# ArXiv Daily Brief - 2026-03-14
|
||
|
||
## 🧠 今日 Top 3(中文可读版)
|
||
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
|
||
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
|
||
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
|
||
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
|
||
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
|
||
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
|
||
- arXiv: http://arxiv.org/abs/2603.12266v1
|
||
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
|
||
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
|
||
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
|
||
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
|
||
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
|
||
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
|
||
- arXiv: http://arxiv.org/abs/2603.12262v1
|
||
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
|
||
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
|
||
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
|
||
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
|
||
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
|
||
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
|
||
- arXiv: http://arxiv.org/abs/2603.12238v1
|
||
|
||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
|
||
- arXiv: http://arxiv.org/abs/2603.12266v1
|
||
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
|
||
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...(cs.CV)
|
||
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
|
||
- arXiv: http://arxiv.org/abs/2603.12262v1
|
||
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
|
||
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously(cs.CV)
|
||
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
|
||
- arXiv: http://arxiv.org/abs/2603.12238v1
|
||
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
|
||
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation(cs.CV)
|
||
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
|
||
- arXiv: http://arxiv.org/abs/2603.12249v1
|
||
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
|
||
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning(cs.CL)
|
||
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
|
||
- arXiv: http://arxiv.org/abs/2603.12246v1
|
||
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
|
||
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training(cs.AI)
|
||
|
||
## 🆕 最新上新 Top 10
|
||
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
|
||
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
|
||
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
|
||
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
|
||
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
|
||
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
|
||
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
|
||
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
|
||
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
|
||
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
|
||
|
||
## Val 今日建议
|
||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|