Files
val-blog/org/cases/arxiv_digest/output/latest.md
T

5.7 KiB
Raw Blame History

ArXiv Daily Brief - 2026-03-14

🧠 今日 Top 3(中文可读版)

  1. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
    • 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
    • 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
    • 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
    • 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
    • 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
    • arXiv: http://arxiv.org/abs/2603.12266v1
  2. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
    • 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
    • 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
    • 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
    • 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
    • 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
    • arXiv: http://arxiv.org/abs/2603.12262v1
  3. SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation
    • 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
    • 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
    • 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
    • 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
    • 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
    • arXiv: http://arxiv.org/abs/2603.12238v1

🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)

  1. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
    • arXiv: http://arxiv.org/abs/2603.12266v1
    • 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
    • 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...cs.CV
  2. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
    • arXiv: http://arxiv.org/abs/2603.12262v1
    • 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
    • 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneouslycs.CV
  3. SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation
    • arXiv: http://arxiv.org/abs/2603.12238v1
    • 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
    • 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generationcs.CV
  4. SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning
    • arXiv: http://arxiv.org/abs/2603.12249v1
    • 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
    • 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoningcs.CL
  5. Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
    • arXiv: http://arxiv.org/abs/2603.12246v1
    • 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
    • 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Trainingcs.AI

🆕 最新上新 Top 10

  1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
  2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
  3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
  4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
  5. Ψ_0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
  6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
  7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
  8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
  9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
  10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1

Val 今日建议

  • 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
  • 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。