5.9 KiB
5.9 KiB
ArXiv Daily Brief - 2026-03-10
🧠 今日 Top 3(中文可读版)
- Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
- 中文题目(意译): Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
- 这篇在讲什么: We challenge the prevailing practice that SOTA VLMs must rely on vision encoders initialized via massive contrastive pre训练 (e.g., CLIP/Si...
- 它怎么做: To address this issue, 提出了 Penguin-VL, whose vision encoder is initialized from a text-only LLM.
- 得出了什么结果: Across various image and video 基准测试s, Penguin-VL achieves performance comparable to leading VLMs (e.g., Qwen3-VL) in mathematical reasoni...
- 可能的影响: This makes it a strong drop-in alternative for compute-efficient VLMs and enables high performance in resource-constrained settings.
- arXiv: http://arxiv.org/abs/2603.06569v1
- Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
- 中文题目(意译): Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
- 这篇在讲什么: However, SOTA approaches exclude critical context through single-pass retrieval, lose data resolution through compression, and exceed LLM...
- 它怎么做: 引入了 Beyond Rows to Reasoning (BRTR), a multimodal agentic framework for spreadsheet understanding that replaces single-pass retrieval wit...
- 得出了什么结果: Supported by over 200 hours of expert human evaluation, BRTR achieves SOTA performance across three frontier spreadsheet understanding 基准...
- 可能的影响: Recent advances in multimodal Retrieval-Augmented Generation (RAG) enable Large Language 模型s (LLMs) to analyze enterprise spreadsheet wor...
- arXiv: http://arxiv.org/abs/2603.06503v1
- Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
- 中文题目(意译): Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
- 这篇在讲什么: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal (
M^2) data to support real-time decision-... - 它怎么做: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal (
M^2) data to support real-time decision-... - 得出了什么结果: Experimental 结果显示 that selectively removing redundant multisource image object labels from cameras with shared fields of view improves de...
- 可能的影响: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal (
M^2) data to support real-time decision-... - arXiv: http://arxiv.org/abs/2603.06544v1
🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
- Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
- arXiv: http://arxiv.org/abs/2603.06569v1
- 类别: cs.CV | HotScore: 22.0 | 作者: Boqiang Zhang, Lei Ke, Ruihan Yang
- 速读: Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders(cs.CV)
- Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
- arXiv: http://arxiv.org/abs/2603.06503v1
- 类别: cs.CL | HotScore: 21.0 | 作者: Anmol Gulati, Sahil Sen, Waqar Sarguroh
- 速读: Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding an...(cs.CL)
- Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
- arXiv: http://arxiv.org/abs/2603.06544v1
- 类别: cs.CV | HotScore: 18.0 | 作者: Yuhan Zhou, Mehri Sattari, Haihua Chen
- 速读: Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving(cs.CV)
- COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
- arXiv: http://arxiv.org/abs/2603.06495v1
- 类别: cs.LG | HotScore: 17.0 | 作者: Kartik Sharma, Rakshit S. Trivedi
- 速读: COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics(cs.LG)
- SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation
- arXiv: http://arxiv.org/abs/2603.06572v1
- 类别: cs.CV | HotScore: 16.0 | 作者: Vishal Thengane, Zhaochong An, Tianjin Huang
- 速读: SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation(cs.CV)
🆕 最新上新 Top 10
- Multimodal Large Language Models as Image Classifiers (cs.CV) - http://arxiv.org/abs/2603.06578v1
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion (cs.CV) - http://arxiv.org/abs/2603.06577v1
- BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations (cs.CV) - http://arxiv.org/abs/2603.06576v1
- Fly360: Omnidirectional Obstacle Avoidance within Drone View (cs.RO) - http://arxiv.org/abs/2603.06573v1
- SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation (cs.CV) - http://arxiv.org/abs/2603.06572v1
- SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning (cs.CV) - http://arxiv.org/abs/2603.06570v1
- Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders (cs.CV) - http://arxiv.org/abs/2603.06569v1
- A recipe for scalable attention-based MLIPs: unlocking long-range accuracy with all-to-all node attention (cs.LG) - http://arxiv.org/abs/2603.06567v1
- Boosting deep Reinforcement Learning using pretraining with Logical Options (cs.AI) - http://arxiv.org/abs/2603.06565v1
- EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking (cs.CV) - http://arxiv.org/abs/2603.06561v1
Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。