chore: allowlist trusted plugins
This commit is contained in:
Vendored
BIN
Binary file not shown.
@@ -26,7 +26,7 @@
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
|
||||
- arXiv: http://arxiv.org/abs/2603.05483v1
|
||||
- 类别: cs.LG | HotScore: 23.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
|
||||
- 类别: cs.LG | HotScore: 21.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
|
||||
- 速读: SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival An...(cs.LG)
|
||||
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
|
||||
- arXiv: http://arxiv.org/abs/2603.05503v1
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# ArXiv Daily Brief - 2026-03-10
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders**
|
||||
- 中文题目(意译): Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
|
||||
- 这篇在讲什么: We challenge the prevailing practice that SOTA VLMs must rely on vision encoders initialized via massive contrastive pre训练 (e.g., CLIP/Si...
|
||||
- 它怎么做: To address this issue, 提出了 Penguin-VL, whose vision encoder is initialized from a text-only LLM.
|
||||
- 得出了什么结果: Across various image and video 基准测试s, Penguin-VL achieves performance comparable to leading VLMs (e.g., Qwen3-VL) in mathematical reasoni...
|
||||
- 可能的影响: This makes it a strong drop-in alternative for compute-efficient VLMs and enables high performance in resource-constrained settings.
|
||||
- arXiv: http://arxiv.org/abs/2603.06569v1
|
||||
2. **Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing**
|
||||
- 中文题目(意译): Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
|
||||
- 这篇在讲什么: However, SOTA approaches exclude critical context through single-pass retrieval, lose data resolution through compression, and exceed LLM...
|
||||
- 它怎么做: 引入了 Beyond Rows to Reasoning (BRTR), a multimodal agentic framework for spreadsheet understanding that replaces single-pass retrieval wit...
|
||||
- 得出了什么结果: Supported by over 200 hours of expert human evaluation, BRTR achieves SOTA performance across three frontier spreadsheet understanding 基准...
|
||||
- 可能的影响: Recent advances in multimodal Retrieval-Augmented Generation (RAG) enable Large Language 模型s (LLMs) to analyze enterprise spreadsheet wor...
|
||||
- arXiv: http://arxiv.org/abs/2603.06503v1
|
||||
3. **Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving**
|
||||
- 中文题目(意译): Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
|
||||
- 这篇在讲什么: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-...
|
||||
- 它怎么做: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-...
|
||||
- 得出了什么结果: Experimental 结果显示 that selectively removing redundant multisource image object labels from cameras with shared fields of view improves de...
|
||||
- 可能的影响: Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-...
|
||||
- arXiv: http://arxiv.org/abs/2603.06544v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders**
|
||||
- arXiv: http://arxiv.org/abs/2603.06569v1
|
||||
- 类别: cs.CV | HotScore: 22.0 | 作者: Boqiang Zhang, Lei Ke, Ruihan Yang
|
||||
- 速读: Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders(cs.CV)
|
||||
2. **Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing**
|
||||
- arXiv: http://arxiv.org/abs/2603.06503v1
|
||||
- 类别: cs.CL | HotScore: 21.0 | 作者: Anmol Gulati, Sahil Sen, Waqar Sarguroh
|
||||
- 速读: Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding an...(cs.CL)
|
||||
3. **Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving**
|
||||
- arXiv: http://arxiv.org/abs/2603.06544v1
|
||||
- 类别: cs.CV | HotScore: 18.0 | 作者: Yuhan Zhou, Mehri Sattari, Haihua Chen
|
||||
- 速读: Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving(cs.CV)
|
||||
4. **COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics**
|
||||
- arXiv: http://arxiv.org/abs/2603.06495v1
|
||||
- 类别: cs.LG | HotScore: 17.0 | 作者: Kartik Sharma, Rakshit S. Trivedi
|
||||
- 速读: COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics(cs.LG)
|
||||
5. **SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation**
|
||||
- arXiv: http://arxiv.org/abs/2603.06572v1
|
||||
- 类别: cs.CV | HotScore: 16.0 | 作者: Vishal Thengane, Zhaochong An, Tianjin Huang
|
||||
- 速读: SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation(cs.CV)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. Multimodal Large Language Models as Image Classifiers (cs.CV) - http://arxiv.org/abs/2603.06578v1
|
||||
2. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion (cs.CV) - http://arxiv.org/abs/2603.06577v1
|
||||
3. BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations (cs.CV) - http://arxiv.org/abs/2603.06576v1
|
||||
4. Fly360: Omnidirectional Obstacle Avoidance within Drone View (cs.RO) - http://arxiv.org/abs/2603.06573v1
|
||||
5. SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation (cs.CV) - http://arxiv.org/abs/2603.06572v1
|
||||
6. SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning (cs.CV) - http://arxiv.org/abs/2603.06570v1
|
||||
7. Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders (cs.CV) - http://arxiv.org/abs/2603.06569v1
|
||||
8. A recipe for scalable attention-based MLIPs: unlocking long-range accuracy with all-to-all node attention (cs.LG) - http://arxiv.org/abs/2603.06567v1
|
||||
9. Boosting deep Reinforcement Learning using pretraining with Logical Options (cs.AI) - http://arxiv.org/abs/2603.06565v1
|
||||
10. EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking (cs.CV) - http://arxiv.org/abs/2603.06561v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|
||||
@@ -0,0 +1,62 @@
|
||||
# ArXiv Daily Brief - 2026-03-11
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
|
||||
- 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成
|
||||
- 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
|
||||
- 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin...
|
||||
- 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41...
|
||||
- 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
|
||||
- arXiv: http://arxiv.org/abs/2603.08652v1
|
||||
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
|
||||
- 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting
|
||||
- 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence.
|
||||
- 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ...
|
||||
- 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
|
||||
- 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
|
||||
- arXiv: http://arxiv.org/abs/2603.08707v1
|
||||
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
|
||||
- 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning
|
||||
- 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
|
||||
- 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
|
||||
- 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o...
|
||||
- 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
|
||||
- arXiv: http://arxiv.org/abs/2603.08655v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
|
||||
- arXiv: http://arxiv.org/abs/2603.08652v1
|
||||
- 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang
|
||||
- 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation(cs.AI)
|
||||
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
|
||||
- arXiv: http://arxiv.org/abs/2603.08707v1
|
||||
- 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith
|
||||
- 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting(cs.LG)
|
||||
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
|
||||
- arXiv: http://arxiv.org/abs/2603.08655v1
|
||||
- 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins
|
||||
- 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning(cs.AI)
|
||||
4. **Agentic Critical Training**
|
||||
- arXiv: http://arxiv.org/abs/2603.08706v1
|
||||
- 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho
|
||||
- 速读: Agentic Critical Training(cs.AI)
|
||||
5. **Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines**
|
||||
- arXiv: http://arxiv.org/abs/2603.08704v1
|
||||
- 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga
|
||||
- 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...(cs.AI)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1
|
||||
2. FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1
|
||||
3. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1
|
||||
4. Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1
|
||||
5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1
|
||||
6. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1
|
||||
7. A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1
|
||||
8. Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1
|
||||
9. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1
|
||||
10. Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|
||||
@@ -0,0 +1,62 @@
|
||||
# ArXiv Daily Brief - 2026-03-12
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems**
|
||||
- 中文题目(意译): MedMASLab: A Unified Orchestration Framework for 基准ing Multimodal 医疗 Multi-Agent Systems
|
||||
- 这篇在讲什么: Current medical MAS research suffers from non-uniform data ingestion pipelines, inconsistent visual-reasoning evaluation, and a lack of c...
|
||||
- 它怎么做: To address these challenges, 提出了 MedMASLab, a unified framework and 基准测试ing platform for multimodal medical multi-agent systems.
|
||||
- 得出了什么结果: Our systematic evaluation reveals a critical domain-specific performance gap: while MAS improves reasoning depth, current architectures e...
|
||||
- 可能的影响: MedMASLab introduces: (1) A standardized multimodal agent communication protocol that enables seamless integration of 11 heterogeneous MA...
|
||||
- arXiv: http://arxiv.org/abs/2603.09909v1
|
||||
2. **PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs**
|
||||
- 中文题目(意译): PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs
|
||||
- 这篇在讲什么: Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonom...
|
||||
- 它怎么做: Inspired by the hierarchical memory process of human pathologists, 提出了 PathMem, a memory-centric multimodal framework for pathology MLLMs.
|
||||
- 得出了什么结果: PathMem achieves SOTA performance across 基准测试s, improving WSI-Bench report generation (12.8% WSI-Precision, 10.1% WSI-Relevance) and open...
|
||||
- 可能的影响: Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonom...
|
||||
- arXiv: http://arxiv.org/abs/2603.09943v1
|
||||
3. **MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning**
|
||||
- 中文题目(意译): MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
|
||||
- 这篇在讲什么: Existing replay-based strategies, such as fixed interleaved replay, accuracy-supervised, and loss-driven scheduling, remain limited: some...
|
||||
- 它怎么做: Motivated by retention dynamics under sequential fine-tuning, 提出了 Memory-Inspired Sampler and Scheduler Replay (MSSR), an experience repl...
|
||||
- 得出了什么结果: Existing replay-based strategies, such as fixed interleaved replay, accuracy-supervised, and loss-driven scheduling, remain limited: some...
|
||||
- 可能的影响: While strong adaptability enables rapid acquisition of new knowledge, it also exposes LLMs to catastrophic forgetting, where previously l...
|
||||
- arXiv: http://arxiv.org/abs/2603.09892v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems**
|
||||
- arXiv: http://arxiv.org/abs/2603.09909v1
|
||||
- 类别: cs.AI | HotScore: 64.96 | 作者: Yunhang Qian, Xiaobin Hu, Jiaquan Yu
|
||||
- 速读: MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-...(cs.AI)
|
||||
2. **PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs**
|
||||
- arXiv: http://arxiv.org/abs/2603.09943v1
|
||||
- 类别: cs.AI | HotScore: 57.41 | 作者: Jinyue Li, Yuci Liang, Qiankun Li
|
||||
- 速读: PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs(cs.AI)
|
||||
3. **MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning**
|
||||
- arXiv: http://arxiv.org/abs/2603.09892v1
|
||||
- 类别: cs.LG | HotScore: 56.77 | 作者: Yiyang Lu, Yu He, Jianlong Chen
|
||||
- 速读: MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning(cs.LG)
|
||||
4. **From Data Statistics to Feature Geometry: How Correlations Shape Superposition**
|
||||
- arXiv: http://arxiv.org/abs/2603.09972v1
|
||||
- 类别: cs.LG | HotScore: 55.74 | 作者: Lucas Prieto, Edward Stevinson, Melih Barsbey
|
||||
- 速读: From Data Statistics to Feature Geometry: How Correlations Shape Superposition(cs.LG)
|
||||
5. **Think Before You Lie: How Reasoning Improves Honesty**
|
||||
- arXiv: http://arxiv.org/abs/2603.09957v1
|
||||
- 类别: cs.AI | HotScore: 55.65 | 作者: Ann Yuan, Asma Ghandeharioun, Carter Blum
|
||||
- 速读: Think Before You Lie: How Reasoning Improves Honesty(cs.AI)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. Task Aware Modulation Using Representation Learning for Upsaling of Terrestrial Carbon Fluxes (cs.LG) - http://arxiv.org/abs/2603.09974v1
|
||||
2. From Data Statistics to Feature Geometry: How Correlations Shape Superposition (cs.LG) - http://arxiv.org/abs/2603.09972v1
|
||||
3. TiPToP: A Modular Open-Vocabulary Planning System for Robotic Manipulation (cs.RO) - http://arxiv.org/abs/2603.09971v1
|
||||
4. CREATE: Testing LLMs for Associative Creativity (cs.CL) - http://arxiv.org/abs/2603.09970v1
|
||||
5. ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare (cs.CV) - http://arxiv.org/abs/2603.09968v1
|
||||
6. Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision People (cs.HC) - http://arxiv.org/abs/2603.09964v1
|
||||
7. Emotional Modulation in Swarm Decision Dynamics (cs.MA) - http://arxiv.org/abs/2603.09963v1
|
||||
8. BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion (cs.RO) - http://arxiv.org/abs/2603.09961v1
|
||||
9. Think Before You Lie: How Reasoning Improves Honesty (cs.AI) - http://arxiv.org/abs/2603.09957v1
|
||||
10. Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization (cs.RO) - http://arxiv.org/abs/2603.09956v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|
||||
@@ -0,0 +1,62 @@
|
||||
# ArXiv Daily Brief - 2026-03-13
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity**
|
||||
- 中文题目(意译): Too Vivid to Be Real? 基准ing and Calibrating Generative Color Fidelity
|
||||
- 这篇在讲什么: Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authent...
|
||||
- 它怎么做: To address this issue, 提出了 Color Fidelity 数据集 (CFD) and Color Fidelity Metric (CFM) for objective evaluation of color fidelity in realist...
|
||||
- 得出了什么结果: Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authent...
|
||||
- 可能的影响: Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authent...
|
||||
- arXiv: http://arxiv.org/abs/2603.10990v1
|
||||
2. **Ranking Reasoning LLMs under Test-Time Scaling**
|
||||
- 中文题目(意译): Ranking Reasoning LLMs under Test-Time Scaling
|
||||
- 这篇在讲什么: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
|
||||
- 它怎么做: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
|
||||
- 得出了什么结果: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
|
||||
- 可能的影响: Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking 模型s in this regime remains underexplored.
|
||||
- arXiv: http://arxiv.org/abs/2603.10960v1
|
||||
3. **Pointy - A Lightweight Transformer for Point Cloud Foundation Models**
|
||||
- 中文题目(意译): Pointy - A Lightweight Transformer for Point Cloud Foundation 模型
|
||||
- 这篇在讲什么: Foundation 模型s for point cloud data have recently grown in capability, often leveraging extensive representation learning from language o...
|
||||
- 它怎么做: Interestingly, 该方法 approaches SOTA results from 模型s that have seen over a million point clouds, images, and text samples, demonstrating t...
|
||||
- 得出了什么结果: In contrast to the heavy reliance on cross-modal supervision, our 模型 is trained only on 39k point clouds - yet it outperforms several lar...
|
||||
- 可能的影响: Foundation 模型s for point cloud data have recently grown in capability, often leveraging extensive representation learning from language o...
|
||||
- arXiv: http://arxiv.org/abs/2603.10963v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity**
|
||||
- arXiv: http://arxiv.org/abs/2603.10990v1
|
||||
- 类别: cs.CV | HotScore: 58.02 | 作者: Zhengyao Fang, Zexi Jia, Yijia Zhong
|
||||
- 速读: Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity(cs.CV)
|
||||
2. **Ranking Reasoning LLMs under Test-Time Scaling**
|
||||
- arXiv: http://arxiv.org/abs/2603.10960v1
|
||||
- 类别: cs.LG | HotScore: 54.59 | 作者: Mohsen Hariri, Michael Hinczewski, Jing Ma
|
||||
- 速读: Ranking Reasoning LLMs under Test-Time Scaling(cs.LG)
|
||||
3. **Pointy - A Lightweight Transformer for Point Cloud Foundation Models**
|
||||
- arXiv: http://arxiv.org/abs/2603.10963v1
|
||||
- 类别: cs.CV | HotScore: 51.64 | 作者: Konrad Szafer, Marek Kraft, Dominik Belter
|
||||
- 速读: Pointy - A Lightweight Transformer for Point Cloud Foundation Models(cs.CV)
|
||||
4. **Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment**
|
||||
- arXiv: http://arxiv.org/abs/2603.10929v1
|
||||
- 类别: cs.CV | HotScore: 51.11 | 作者: Fanqi Yu, Matteo Tiezzi, Tommaso Apicella
|
||||
- 速读: Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment(cs.CV)
|
||||
5. **Bio-Inspired Self-Supervised Learning for Wrist-worn IMU Signals**
|
||||
- arXiv: http://arxiv.org/abs/2603.10961v1
|
||||
- 类别: cs.LG | HotScore: 50.61 | 作者: Prithviraj Tarale, Kiet Chu, Abhishek Varghese
|
||||
- 速读: Bio-Inspired Self-Supervised Learning for Wrist-worn IMU Signals(cs.LG)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. COMIC: Agentic Sketch Comedy Generation (cs.CV) - http://arxiv.org/abs/2603.11048v1
|
||||
2. LiTo: Surface Light Field Tokenization (cs.CV) - http://arxiv.org/abs/2603.11047v1
|
||||
3. Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation (cs.LG) - http://arxiv.org/abs/2603.11045v1
|
||||
4. Agentar-Fin-OCR (cs.CV) - http://arxiv.org/abs/2603.11044v1
|
||||
5. V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation (cs.CV) - http://arxiv.org/abs/2603.11042v1
|
||||
6. DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving (cs.CV) - http://arxiv.org/abs/2603.11041v1
|
||||
7. Instruction set for the representation of graphs (cs.CL) - http://arxiv.org/abs/2603.11039v1
|
||||
8. Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge (cs.CL) - http://arxiv.org/abs/2603.11027v1
|
||||
9. Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style (cs.CV) - http://arxiv.org/abs/2603.11024v1
|
||||
10. Leech Lattice Vector Quantization for Efficient LLM Compression (cs.LG) - http://arxiv.org/abs/2603.11021v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|
||||
@@ -0,0 +1,62 @@
|
||||
# ArXiv Daily Brief - 2026-03-14
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
|
||||
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
|
||||
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
|
||||
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
|
||||
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
|
||||
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
|
||||
- arXiv: http://arxiv.org/abs/2603.12266v1
|
||||
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
|
||||
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
|
||||
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
|
||||
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
|
||||
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
|
||||
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
|
||||
- arXiv: http://arxiv.org/abs/2603.12262v1
|
||||
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
|
||||
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
|
||||
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
|
||||
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
|
||||
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
|
||||
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
|
||||
- arXiv: http://arxiv.org/abs/2603.12238v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
|
||||
- arXiv: http://arxiv.org/abs/2603.12266v1
|
||||
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
|
||||
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...(cs.CV)
|
||||
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
|
||||
- arXiv: http://arxiv.org/abs/2603.12262v1
|
||||
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
|
||||
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously(cs.CV)
|
||||
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
|
||||
- arXiv: http://arxiv.org/abs/2603.12238v1
|
||||
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
|
||||
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation(cs.CV)
|
||||
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
|
||||
- arXiv: http://arxiv.org/abs/2603.12249v1
|
||||
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
|
||||
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning(cs.CL)
|
||||
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
|
||||
- arXiv: http://arxiv.org/abs/2603.12246v1
|
||||
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
|
||||
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training(cs.AI)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
|
||||
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
|
||||
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
|
||||
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
|
||||
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
|
||||
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
|
||||
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
|
||||
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
|
||||
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
|
||||
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|
||||
@@ -1,61 +1,61 @@
|
||||
# ArXiv Daily Brief - 2026-03-09
|
||||
# ArXiv Daily Brief - 2026-03-14
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
|
||||
- 中文题目(意译): SurvHTE-Bench: A 基准 for Heterogeneous Treatment Effect Estimation in Survival Analysis
|
||||
- 这篇在讲什么: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
|
||||
- 它怎么做: 引入了 SurvHTE-Bench, the first comprehensive 基准测试 for HTE estimation with censored outcomes.
|
||||
- 得出了什么结果: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
|
||||
- 可能的影响: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
|
||||
- arXiv: http://arxiv.org/abs/2603.05483v1
|
||||
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
|
||||
- 中文题目(意译): 加速 Text-to-Video 生成 with Calibrated Sparse Attention
|
||||
- 这篇在讲什么: Recent 扩散 模型s enable high-quality 视频生成, but suffer from slow runtimes.
|
||||
- 它怎么做: Motivated by this, 引入了 CalibAtt, a 训练-free method that accelerates 视频生成 via calibrated sparse attention.
|
||||
- 得出了什么结果: Extensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled 模型s at various resolutions show that CalibAtt achieves up to 1.58x ...
|
||||
- 可能的影响: Recent 扩散 模型s enable high-quality 视频生成, but suffer from slow runtimes.
|
||||
- arXiv: http://arxiv.org/abs/2603.05503v1
|
||||
3. **Observing and Controlling Features in Vision-Language-Action Models**
|
||||
- 中文题目(意译): Observing and 控制 特征 in 视觉-语言-动作 模型
|
||||
- 这篇在讲什么: 视觉-语言-动作 模型s (VLAs) have shown remarkable progress towards embodied intelligence.
|
||||
- 它怎么做: In this work, 提出了 to close this gap by introducing and analyzing two main concepts: feature-observability and feature-controllability.
|
||||
- 得出了什么结果: Our 结果显示 that targeted, lightweight interventions can reliably steer a robot's behavior while preserving closed-loop capabilities.
|
||||
- 可能的影响: 视觉-语言-动作 模型s (VLAs) have shown remarkable progress towards embodied intelligence.
|
||||
- arXiv: http://arxiv.org/abs/2603.05487v1
|
||||
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
|
||||
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
|
||||
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
|
||||
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
|
||||
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
|
||||
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
|
||||
- arXiv: http://arxiv.org/abs/2603.12266v1
|
||||
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
|
||||
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
|
||||
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
|
||||
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
|
||||
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
|
||||
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
|
||||
- arXiv: http://arxiv.org/abs/2603.12262v1
|
||||
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
|
||||
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
|
||||
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
|
||||
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
|
||||
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
|
||||
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
|
||||
- arXiv: http://arxiv.org/abs/2603.12238v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
|
||||
- arXiv: http://arxiv.org/abs/2603.05483v1
|
||||
- 类别: cs.LG | HotScore: 23.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
|
||||
- 速读: SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival An...(cs.LG)
|
||||
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
|
||||
- arXiv: http://arxiv.org/abs/2603.05503v1
|
||||
- 类别: cs.CV | HotScore: 18.0 | 作者: Shai Yehezkel, Shahar Yadin, Noam Elata
|
||||
- 速读: Accelerating Text-to-Video Generation with Calibrated Sparse Attention(cs.CV)
|
||||
3. **Observing and Controlling Features in Vision-Language-Action Models**
|
||||
- arXiv: http://arxiv.org/abs/2603.05487v1
|
||||
- 类别: cs.RO | HotScore: 18.0 | 作者: Hugo Buurmeijer, Carmen Amo Alonso, Aiden Swann
|
||||
- 速读: Observing and Controlling Features in Vision-Language-Action Models(cs.RO)
|
||||
4. **Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation**
|
||||
- arXiv: http://arxiv.org/abs/2603.05485v1
|
||||
- 类别: cs.AI | HotScore: 17.0 | 作者: Benjamin Feuer, Lucas Rosenblatt, Oussama Elachqar
|
||||
- 速读: Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation(cs.AI)
|
||||
5. **An interpretable prototype parts-based neural network for medical tabular data**
|
||||
- arXiv: http://arxiv.org/abs/2603.05423v1
|
||||
- 类别: cs.LG | HotScore: 17.0 | 作者: Jacek Karolczak, Jerzy Stefanowski
|
||||
- 速读: An interpretable prototype parts-based neural network for medical tabular data(cs.LG)
|
||||
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
|
||||
- arXiv: http://arxiv.org/abs/2603.12266v1
|
||||
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
|
||||
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...(cs.CV)
|
||||
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
|
||||
- arXiv: http://arxiv.org/abs/2603.12262v1
|
||||
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
|
||||
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously(cs.CV)
|
||||
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
|
||||
- arXiv: http://arxiv.org/abs/2603.12238v1
|
||||
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
|
||||
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation(cs.CV)
|
||||
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
|
||||
- arXiv: http://arxiv.org/abs/2603.12249v1
|
||||
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
|
||||
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning(cs.CL)
|
||||
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
|
||||
- arXiv: http://arxiv.org/abs/2603.12246v1
|
||||
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
|
||||
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training(cs.AI)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups (cs.CV) - http://arxiv.org/abs/2603.05507v1
|
||||
2. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning (cs.CV) - http://arxiv.org/abs/2603.05506v1
|
||||
3. RoboPocket: Improve Robot Policies Instantly with Your Phone (cs.RO) - http://arxiv.org/abs/2603.05504v1
|
||||
4. Accelerating Text-to-Video Generation with Calibrated Sparse Attention (cs.CV) - http://arxiv.org/abs/2603.05503v1
|
||||
5. POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation (cs.LG) - http://arxiv.org/abs/2603.05500v1
|
||||
6. The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks (cs.AI) - http://arxiv.org/abs/2603.05498v1
|
||||
7. Safe-SAGE: Social-Semantic Adaptive Guidance for Safe Engagement through Laplace-Modulated Poisson Safety Functions (cs.RO) - http://arxiv.org/abs/2603.05497v1
|
||||
8. Cheap Thrills: Effective Amortized Optimization Using Inexpensive Labels (cs.LG) - http://arxiv.org/abs/2603.05495v1
|
||||
9. Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation (cs.LG) - http://arxiv.org/abs/2603.05494v1
|
||||
10. cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots (cs.RO) - http://arxiv.org/abs/2603.05493v1
|
||||
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
|
||||
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
|
||||
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
|
||||
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
|
||||
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
|
||||
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
|
||||
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
|
||||
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
|
||||
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
|
||||
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
|
||||
@@ -1 +1 @@
|
||||
{ "date": "2026-03-09", "pushed": true, "channel": "gmail" }
|
||||
{"date":"2026-03-14","pushed":true,"channel":"gmail","message_id":"19ce9c43f2d5e7c0","sent_at":"2026-03-14T08:35:00+08:00","cancelled":true,"cancelled_at":"2026-03-14T10:46:00+08:00","cancelled_by":"谷老板指令"}
|
||||
@@ -1,32 +1,32 @@
|
||||
{
|
||||
"updatedAt": "2026-03-09T06:53:12.730807",
|
||||
"date": "2026-03-09",
|
||||
"updatedAt": "2026-03-14T08:34:25.638154",
|
||||
"date": "2026-03-14",
|
||||
"papersFetched": 120,
|
||||
"topHot": [
|
||||
{
|
||||
"title": "SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis",
|
||||
"arxiv_id": "2603.05483v1",
|
||||
"score": 23.0
|
||||
"title": "MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning",
|
||||
"arxiv_id": "2603.12266v1",
|
||||
"score": 56.54
|
||||
},
|
||||
{
|
||||
"title": "Accelerating Text-to-Video Generation with Calibrated Sparse Attention",
|
||||
"arxiv_id": "2603.05503v1",
|
||||
"score": 18.0
|
||||
"title": "Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously",
|
||||
"arxiv_id": "2603.12262v1",
|
||||
"score": 56.53
|
||||
},
|
||||
{
|
||||
"title": "Observing and Controlling Features in Vision-Language-Action Models",
|
||||
"arxiv_id": "2603.05487v1",
|
||||
"score": 18.0
|
||||
"title": "SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation",
|
||||
"arxiv_id": "2603.12238v1",
|
||||
"score": 56.46
|
||||
},
|
||||
{
|
||||
"title": "Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation",
|
||||
"arxiv_id": "2603.05485v1",
|
||||
"score": 17.0
|
||||
"title": "SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning",
|
||||
"arxiv_id": "2603.12249v1",
|
||||
"score": 55.5
|
||||
},
|
||||
{
|
||||
"title": "An interpretable prototype parts-based neural network for medical tabular data",
|
||||
"arxiv_id": "2603.05423v1",
|
||||
"score": 17.0
|
||||
"title": "Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training",
|
||||
"arxiv_id": "2603.12246v1",
|
||||
"score": 53.49
|
||||
}
|
||||
]
|
||||
}
|
||||
Reference in New Issue
Block a user