chore: allowlist trusted plugins

This commit is contained in:
Chen Gu
2026-08-13 17:00:21 +08:00
committed by Chen Gu
parent 8027f1186d
commit 319e8a7ec9
66 changed files with 6229 additions and 124 deletions
Binary file not shown.
+1 -1
View File
@@ -26,7 +26,7 @@
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
- arXiv: http://arxiv.org/abs/2603.05483v1
- 类别: cs.LG | HotScore: 23.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
- 类别: cs.LG | HotScore: 21.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
- 速读: SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival An...cs.LG
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
- arXiv: http://arxiv.org/abs/2603.05503v1
@@ -0,0 +1,62 @@
# ArXiv Daily Brief - 2026-03-10
## 🧠 今日 Top 3(中文可读版)
1. **Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders**
- 中文题目(意译): Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
- 这篇在讲什么: We challenge the prevailing practice that SOTA VLMs must rely on vision encoders initialized via massive contrastive pre训练 (e.g., CLIP/Si...
- 它怎么做: To address this issue, 提出了 Penguin-VL, whose vision encoder is initialized from a text-only LLM.
- 得出了什么结果: Across various image and video 基准测试s, Penguin-VL achieves performance comparable to leading VLMs (e.g., Qwen3-VL) in mathematical reasoni...
- 可能的影响: This makes it a strong drop-in alternative for compute-efficient VLMs and enables high performance in resource-constrained settings.
- arXiv: http://arxiv.org/abs/2603.06569v1
2. **Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing**
- 中文题目(意译): Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
- 这篇在讲什么: However, SOTA approaches exclude critical context through single-pass retrieval, lose data resolution through compression, and exceed LLM...
- 它怎么做: 引入了 Beyond Rows to Reasoning (BRTR), a multimodal agentic framework for spreadsheet understanding that replaces single-pass retrieval wit...
- 得出了什么结果: Supported by over 200 hours of expert human evaluation, BRTR achieves SOTA performance across three frontier spreadsheet understanding 基准...
- 可能的影响: Recent advances in multimodal Retrieval-Augmented Generation (RAG) enable Large Language 模型s (LLMs) to analyze enterprise spreadsheet wor...
- arXiv: http://arxiv.org/abs/2603.06503v1
3. **Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving**
- 中文题目(意译): Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
- 这篇在讲什么: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-...
- 它怎么做: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-...
- 得出了什么结果: Experimental 结果显示 that selectively removing redundant multisource image object labels from cameras with shared fields of view improves de...
- 可能的影响: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-...
- arXiv: http://arxiv.org/abs/2603.06544v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders**
- arXiv: http://arxiv.org/abs/2603.06569v1
- 类别: cs.CV | HotScore: 22.0 | 作者: Boqiang Zhang, Lei Ke, Ruihan Yang
- 速读: Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoderscs.CV
2. **Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing**
- arXiv: http://arxiv.org/abs/2603.06503v1
- 类别: cs.CL | HotScore: 21.0 | 作者: Anmol Gulati, Sahil Sen, Waqar Sarguroh
- 速读: Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding an...cs.CL
3. **Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving**
- arXiv: http://arxiv.org/abs/2603.06544v1
- 类别: cs.CV | HotScore: 18.0 | 作者: Yuhan Zhou, Mehri Sattari, Haihua Chen
- 速读: Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Drivingcs.CV
4. **COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics**
- arXiv: http://arxiv.org/abs/2603.06495v1
- 类别: cs.LG | HotScore: 17.0 | 作者: Kartik Sharma, Rakshit S. Trivedi
- 速读: COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamicscs.LG
5. **SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation**
- arXiv: http://arxiv.org/abs/2603.06572v1
- 类别: cs.CV | HotScore: 16.0 | 作者: Vishal Thengane, Zhaochong An, Tianjin Huang
- 速读: SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentationcs.CV
## 🆕 最新上新 Top 10
1. Multimodal Large Language Models as Image Classifiers (cs.CV) - http://arxiv.org/abs/2603.06578v1
2. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion (cs.CV) - http://arxiv.org/abs/2603.06577v1
3. BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations (cs.CV) - http://arxiv.org/abs/2603.06576v1
4. Fly360: Omnidirectional Obstacle Avoidance within Drone View (cs.RO) - http://arxiv.org/abs/2603.06573v1
5. SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation (cs.CV) - http://arxiv.org/abs/2603.06572v1
6. SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning (cs.CV) - http://arxiv.org/abs/2603.06570v1
7. Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders (cs.CV) - http://arxiv.org/abs/2603.06569v1
8. A recipe for scalable attention-based MLIPs: unlocking long-range accuracy with all-to-all node attention (cs.LG) - http://arxiv.org/abs/2603.06567v1
9. Boosting deep Reinforcement Learning using pretraining with Logical Options (cs.AI) - http://arxiv.org/abs/2603.06565v1
10. EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking (cs.CV) - http://arxiv.org/abs/2603.06561v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
@@ -0,0 +1,62 @@
# ArXiv Daily Brief - 2026-03-11
## 🧠 今日 Top 3(中文可读版)
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
- 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成
- 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
- 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin...
- 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41...
- 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
- arXiv: http://arxiv.org/abs/2603.08652v1
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
- 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting
- 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence.
- 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ...
- 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
- 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
- arXiv: http://arxiv.org/abs/2603.08707v1
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
- 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning
- 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o...
- 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- arXiv: http://arxiv.org/abs/2603.08655v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
- arXiv: http://arxiv.org/abs/2603.08652v1
- 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang
- 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generationcs.AI
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
- arXiv: http://arxiv.org/abs/2603.08707v1
- 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith
- 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecastingcs.LG
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
- arXiv: http://arxiv.org/abs/2603.08655v1
- 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins
- 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoningcs.AI
4. **Agentic Critical Training**
- arXiv: http://arxiv.org/abs/2603.08706v1
- 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho
- 速读: Agentic Critical Trainingcs.AI
5. **Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines**
- arXiv: http://arxiv.org/abs/2603.08704v1
- 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga
- 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...cs.AI
## 🆕 最新上新 Top 10
1. Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1
2. FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1
3. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1
4. Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1
5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1
6. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1
7. A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1
8. Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1
9. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1
10. Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
@@ -0,0 +1,62 @@
# ArXiv Daily Brief - 2026-03-12
## 🧠 今日 Top 3(中文可读版)
1. **MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems**
- 中文题目(意译): MedMASLab: A Unified Orchestration Framework for 基准ing Multimodal 医疗 Multi-Agent Systems
- 这篇在讲什么: Current medical MAS research suffers from non-uniform data ingestion pipelines, inconsistent visual-reasoning evaluation, and a lack of c...
- 它怎么做: To address these challenges, 提出了 MedMASLab, a unified framework and 基准测试ing platform for multimodal medical multi-agent systems.
- 得出了什么结果: Our systematic evaluation reveals a critical domain-specific performance gap: while MAS improves reasoning depth, current architectures e...
- 可能的影响: MedMASLab introduces: (1) A standardized multimodal agent communication protocol that enables seamless integration of 11 heterogeneous MA...
- arXiv: http://arxiv.org/abs/2603.09909v1
2. **PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs**
- 中文题目(意译): PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs
- 这篇在讲什么: Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonom...
- 它怎么做: Inspired by the hierarchical memory process of human pathologists, 提出了 PathMem, a memory-centric multimodal framework for pathology MLLMs.
- 得出了什么结果: PathMem achieves SOTA performance across 基准测试s, improving WSI-Bench report generation (12.8% WSI-Precision, 10.1% WSI-Relevance) and open...
- 可能的影响: Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonom...
- arXiv: http://arxiv.org/abs/2603.09943v1
3. **MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning**
- 中文题目(意译): MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
- 这篇在讲什么: Existing replay-based strategies, such as fixed interleaved replay, accuracy-supervised, and loss-driven scheduling, remain limited: some...
- 它怎么做: Motivated by retention dynamics under sequential fine-tuning, 提出了 Memory-Inspired Sampler and Scheduler Replay (MSSR), an experience repl...
- 得出了什么结果: Existing replay-based strategies, such as fixed interleaved replay, accuracy-supervised, and loss-driven scheduling, remain limited: some...
- 可能的影响: While strong adaptability enables rapid acquisition of new knowledge, it also exposes LLMs to catastrophic forgetting, where previously l...
- arXiv: http://arxiv.org/abs/2603.09892v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems**
- arXiv: http://arxiv.org/abs/2603.09909v1
- 类别: cs.AI | HotScore: 64.96 | 作者: Yunhang Qian, Xiaobin Hu, Jiaquan Yu
- 速读: MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-...cs.AI
2. **PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs**
- arXiv: http://arxiv.org/abs/2603.09943v1
- 类别: cs.AI | HotScore: 57.41 | 作者: Jinyue Li, Yuci Liang, Qiankun Li
- 速读: PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMscs.AI
3. **MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning**
- arXiv: http://arxiv.org/abs/2603.09892v1
- 类别: cs.LG | HotScore: 56.77 | 作者: Yiyang Lu, Yu He, Jianlong Chen
- 速读: MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuningcs.LG
4. **From Data Statistics to Feature Geometry: How Correlations Shape Superposition**
- arXiv: http://arxiv.org/abs/2603.09972v1
- 类别: cs.LG | HotScore: 55.74 | 作者: Lucas Prieto, Edward Stevinson, Melih Barsbey
- 速读: From Data Statistics to Feature Geometry: How Correlations Shape Superpositioncs.LG
5. **Think Before You Lie: How Reasoning Improves Honesty**
- arXiv: http://arxiv.org/abs/2603.09957v1
- 类别: cs.AI | HotScore: 55.65 | 作者: Ann Yuan, Asma Ghandeharioun, Carter Blum
- 速读: Think Before You Lie: How Reasoning Improves Honestycs.AI
## 🆕 最新上新 Top 10
1. Task Aware Modulation Using Representation Learning for Upsaling of Terrestrial Carbon Fluxes (cs.LG) - http://arxiv.org/abs/2603.09974v1
2. From Data Statistics to Feature Geometry: How Correlations Shape Superposition (cs.LG) - http://arxiv.org/abs/2603.09972v1
3. TiPToP: A Modular Open-Vocabulary Planning System for Robotic Manipulation (cs.RO) - http://arxiv.org/abs/2603.09971v1
4. CREATE: Testing LLMs for Associative Creativity (cs.CL) - http://arxiv.org/abs/2603.09970v1
5. ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare (cs.CV) - http://arxiv.org/abs/2603.09968v1
6. Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision People (cs.HC) - http://arxiv.org/abs/2603.09964v1
7. Emotional Modulation in Swarm Decision Dynamics (cs.MA) - http://arxiv.org/abs/2603.09963v1
8. BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion (cs.RO) - http://arxiv.org/abs/2603.09961v1
9. Think Before You Lie: How Reasoning Improves Honesty (cs.AI) - http://arxiv.org/abs/2603.09957v1
10. Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization (cs.RO) - http://arxiv.org/abs/2603.09956v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
@@ -0,0 +1,62 @@
# ArXiv Daily Brief - 2026-03-13
## 🧠 今日 Top 3(中文可读版)
1. **Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity**
- 中文题目(意译): Too Vivid to Be Real? 基准ing and Calibrating Generative Color Fidelity
- 这篇在讲什么: Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authent...
- 它怎么做: To address this issue, 提出了 Color Fidelity 数据集 (CFD) and Color Fidelity Metric (CFM) for objective evaluation of color fidelity in realist...
- 得出了什么结果: Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authent...
- 可能的影响: Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authent...
- arXiv: http://arxiv.org/abs/2603.10990v1
2. **Ranking Reasoning LLMs under Test-Time Scaling**
- 中文题目(意译): Ranking Reasoning LLMs under Test-Time Scaling
- 这篇在讲什么: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
- 它怎么做: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
- 得出了什么结果: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
- 可能的影响: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
- arXiv: http://arxiv.org/abs/2603.10960v1
3. **Pointy - A Lightweight Transformer for Point Cloud Foundation Models**
- 中文题目(意译): Pointy - A Lightweight Transformer for Point Cloud Foundation 模型
- 这篇在讲什么: Foundation 模型s for point cloud data have recently grown in capability, often leveraging extensive representation learning from language o...
- 它怎么做: Interestingly, 该方法 approaches SOTA results from 模型s that have seen over a million point clouds, images, and text samples, demonstrating t...
- 得出了什么结果: In contrast to the heavy reliance on cross-modal supervision, our 模型 is trained only on 39k point clouds - yet it outperforms several lar...
- 可能的影响: Foundation 模型s for point cloud data have recently grown in capability, often leveraging extensive representation learning from language o...
- arXiv: http://arxiv.org/abs/2603.10963v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity**
- arXiv: http://arxiv.org/abs/2603.10990v1
- 类别: cs.CV | HotScore: 58.02 | 作者: Zhengyao Fang, Zexi Jia, Yijia Zhong
- 速读: Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelitycs.CV
2. **Ranking Reasoning LLMs under Test-Time Scaling**
- arXiv: http://arxiv.org/abs/2603.10960v1
- 类别: cs.LG | HotScore: 54.59 | 作者: Mohsen Hariri, Michael Hinczewski, Jing Ma
- 速读: Ranking Reasoning LLMs under Test-Time Scalingcs.LG
3. **Pointy - A Lightweight Transformer for Point Cloud Foundation Models**
- arXiv: http://arxiv.org/abs/2603.10963v1
- 类别: cs.CV | HotScore: 51.64 | 作者: Konrad Szafer, Marek Kraft, Dominik Belter
- 速读: Pointy - A Lightweight Transformer for Point Cloud Foundation Modelscs.CV
4. **Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment**
- arXiv: http://arxiv.org/abs/2603.10929v1
- 类别: cs.CV | HotScore: 51.11 | 作者: Fanqi Yu, Matteo Tiezzi, Tommaso Apicella
- 速读: Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustmentcs.CV
5. **Bio-Inspired Self-Supervised Learning for Wrist-worn IMU Signals**
- arXiv: http://arxiv.org/abs/2603.10961v1
- 类别: cs.LG | HotScore: 50.61 | 作者: Prithviraj Tarale, Kiet Chu, Abhishek Varghese
- 速读: Bio-Inspired Self-Supervised Learning for Wrist-worn IMU Signalscs.LG
## 🆕 最新上新 Top 10
1. COMIC: Agentic Sketch Comedy Generation (cs.CV) - http://arxiv.org/abs/2603.11048v1
2. LiTo: Surface Light Field Tokenization (cs.CV) - http://arxiv.org/abs/2603.11047v1
3. Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation (cs.LG) - http://arxiv.org/abs/2603.11045v1
4. Agentar-Fin-OCR (cs.CV) - http://arxiv.org/abs/2603.11044v1
5. V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation (cs.CV) - http://arxiv.org/abs/2603.11042v1
6. DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving (cs.CV) - http://arxiv.org/abs/2603.11041v1
7. Instruction set for the representation of graphs (cs.CL) - http://arxiv.org/abs/2603.11039v1
8. Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge (cs.CL) - http://arxiv.org/abs/2603.11027v1
9. Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style (cs.CV) - http://arxiv.org/abs/2603.11024v1
10. Leech Lattice Vector Quantization for Efficient LLM Compression (cs.LG) - http://arxiv.org/abs/2603.11021v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
@@ -0,0 +1,62 @@
# ArXiv Daily Brief - 2026-03-14
## 🧠 今日 Top 3(中文可读版)
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- arXiv: http://arxiv.org/abs/2603.12266v1
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- arXiv: http://arxiv.org/abs/2603.12262v1
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- arXiv: http://arxiv.org/abs/2603.12238v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- arXiv: http://arxiv.org/abs/2603.12266v1
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...cs.CV
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- arXiv: http://arxiv.org/abs/2603.12262v1
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneouslycs.CV
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- arXiv: http://arxiv.org/abs/2603.12238v1
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generationcs.CV
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
- arXiv: http://arxiv.org/abs/2603.12249v1
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoningcs.CL
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
- arXiv: http://arxiv.org/abs/2603.12246v1
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Trainingcs.AI
## 🆕 最新上新 Top 10
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
+52 -52
View File
@@ -1,61 +1,61 @@
# ArXiv Daily Brief - 2026-03-09
# ArXiv Daily Brief - 2026-03-14
## 🧠 今日 Top 3(中文可读版)
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
- 中文题目(意译): SurvHTE-Bench: A 基准 for Heterogeneous Treatment Effect Estimation in Survival Analysis
- 这篇在讲什么: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
- 它怎么做: 引入了 SurvHTE-Bench, the first comprehensive 基准测试 for HTE estimation with censored outcomes.
- 得出了什么结果: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
- 可能的影响: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
- arXiv: http://arxiv.org/abs/2603.05483v1
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
- 中文题目(意译): 加速 Text-to-Video 生成 with Calibrated Sparse Attention
- 这篇在讲什么: Recent 扩散 模型s enable high-quality 视频生成, but suffer from slow runtimes.
- 它怎么做: Motivated by this, 引入了 CalibAtt, a 训练-free method that accelerates 视频生成 via calibrated sparse attention.
- 得出了什么结果: Extensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled 模型s at various resolutions show that CalibAtt achieves up to 1.58x ...
- 可能的影响: Recent 扩散 模型s enable high-quality 视频生成, but suffer from slow runtimes.
- arXiv: http://arxiv.org/abs/2603.05503v1
3. **Observing and Controlling Features in Vision-Language-Action Models**
- 中文题目(意译): Observing and 控制 特征 in 视觉-语言-动作 模型
- 这篇在讲什么: 视觉-语言-动作 模型s (VLAs) have shown remarkable progress towards embodied intelligence.
- 它怎么做: In this work, 提出了 to close this gap by introducing and analyzing two main concepts: feature-observability and feature-controllability.
- 得出了什么结果: Our 结果显示 that targeted, lightweight interventions can reliably steer a robot's behavior while preserving closed-loop capabilities.
- 可能的影响: 视觉-语言-动作 模型s (VLAs) have shown remarkable progress towards embodied intelligence.
- arXiv: http://arxiv.org/abs/2603.05487v1
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- arXiv: http://arxiv.org/abs/2603.12266v1
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- arXiv: http://arxiv.org/abs/2603.12262v1
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- arXiv: http://arxiv.org/abs/2603.12238v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
- arXiv: http://arxiv.org/abs/2603.05483v1
- 类别: cs.LG | HotScore: 23.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
- 速读: SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival An...cs.LG
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
- arXiv: http://arxiv.org/abs/2603.05503v1
- 类别: cs.CV | HotScore: 18.0 | 作者: Shai Yehezkel, Shahar Yadin, Noam Elata
- 速读: Accelerating Text-to-Video Generation with Calibrated Sparse Attentioncs.CV
3. **Observing and Controlling Features in Vision-Language-Action Models**
- arXiv: http://arxiv.org/abs/2603.05487v1
- 类别: cs.RO | HotScore: 18.0 | 作者: Hugo Buurmeijer, Carmen Amo Alonso, Aiden Swann
- 速读: Observing and Controlling Features in Vision-Language-Action Modelscs.RO
4. **Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation**
- arXiv: http://arxiv.org/abs/2603.05485v1
- 类别: cs.AI | HotScore: 17.0 | 作者: Benjamin Feuer, Lucas Rosenblatt, Oussama Elachqar
- 速读: Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluationcs.AI
5. **An interpretable prototype parts-based neural network for medical tabular data**
- arXiv: http://arxiv.org/abs/2603.05423v1
- 类别: cs.LG | HotScore: 17.0 | 作者: Jacek Karolczak, Jerzy Stefanowski
- 速读: An interpretable prototype parts-based neural network for medical tabular datacs.LG
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- arXiv: http://arxiv.org/abs/2603.12266v1
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...cs.CV
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- arXiv: http://arxiv.org/abs/2603.12262v1
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneouslycs.CV
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- arXiv: http://arxiv.org/abs/2603.12238v1
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generationcs.CV
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
- arXiv: http://arxiv.org/abs/2603.12249v1
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoningcs.CL
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
- arXiv: http://arxiv.org/abs/2603.12246v1
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Trainingcs.AI
## 🆕 最新上新 Top 10
1. Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups (cs.CV) - http://arxiv.org/abs/2603.05507v1
2. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning (cs.CV) - http://arxiv.org/abs/2603.05506v1
3. RoboPocket: Improve Robot Policies Instantly with Your Phone (cs.RO) - http://arxiv.org/abs/2603.05504v1
4. Accelerating Text-to-Video Generation with Calibrated Sparse Attention (cs.CV) - http://arxiv.org/abs/2603.05503v1
5. POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation (cs.LG) - http://arxiv.org/abs/2603.05500v1
6. The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks (cs.AI) - http://arxiv.org/abs/2603.05498v1
7. Safe-SAGE: Social-Semantic Adaptive Guidance for Safe Engagement through Laplace-Modulated Poisson Safety Functions (cs.RO) - http://arxiv.org/abs/2603.05497v1
8. Cheap Thrills: Effective Amortized Optimization Using Inexpensive Labels (cs.LG) - http://arxiv.org/abs/2603.05495v1
9. Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation (cs.LG) - http://arxiv.org/abs/2603.05494v1
10. cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots (cs.RO) - http://arxiv.org/abs/2603.05493v1
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
+1 -1
View File
@@ -1 +1 @@
{ "date": "2026-03-09", "pushed": true, "channel": "gmail" }
{"date":"2026-03-14","pushed":true,"channel":"gmail","message_id":"19ce9c43f2d5e7c0","sent_at":"2026-03-14T08:35:00+08:00","cancelled":true,"cancelled_at":"2026-03-14T10:46:00+08:00","cancelled_by":"谷老板指令"}
+17 -17
View File
@@ -1,32 +1,32 @@
{
"updatedAt": "2026-03-09T06:53:12.730807",
"date": "2026-03-09",
"updatedAt": "2026-03-14T08:34:25.638154",
"date": "2026-03-14",
"papersFetched": 120,
"topHot": [
{
"title": "SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis",
"arxiv_id": "2603.05483v1",
"score": 23.0
"title": "MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning",
"arxiv_id": "2603.12266v1",
"score": 56.54
},
{
"title": "Accelerating Text-to-Video Generation with Calibrated Sparse Attention",
"arxiv_id": "2603.05503v1",
"score": 18.0
"title": "Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously",
"arxiv_id": "2603.12262v1",
"score": 56.53
},
{
"title": "Observing and Controlling Features in Vision-Language-Action Models",
"arxiv_id": "2603.05487v1",
"score": 18.0
"title": "SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation",
"arxiv_id": "2603.12238v1",
"score": 56.46
},
{
"title": "Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation",
"arxiv_id": "2603.05485v1",
"score": 17.0
"title": "SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning",
"arxiv_id": "2603.12249v1",
"score": 55.5
},
{
"title": "An interpretable prototype parts-based neural network for medical tabular data",
"arxiv_id": "2603.05423v1",
"score": 17.0
"title": "Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training",
"arxiv_id": "2603.12246v1",
"score": 53.49
}
]
}