chore: allowlist trusted plugins

This commit is contained in:
Chen Gu
2026-08-13 17:00:21 +08:00
committed by Chen Gu
parent 8027f1186d
commit 319e8a7ec9
66 changed files with 6229 additions and 124 deletions
+52 -52
View File
@@ -1,61 +1,61 @@
# ArXiv Daily Brief - 2026-03-09
# ArXiv Daily Brief - 2026-03-14
## 🧠 今日 Top 3(中文可读版)
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
- 中文题目(意译): SurvHTE-Bench: A 基准 for Heterogeneous Treatment Effect Estimation in Survival Analysis
- 这篇在讲什么: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
- 它怎么做: 引入了 SurvHTE-Bench, the first comprehensive 基准测试 for HTE estimation with censored outcomes.
- 得出了什么结果: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
- 可能的影响: Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as preci...
- arXiv: http://arxiv.org/abs/2603.05483v1
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
- 中文题目(意译): 加速 Text-to-Video 生成 with Calibrated Sparse Attention
- 这篇在讲什么: Recent 扩散 模型s enable high-quality 视频生成, but suffer from slow runtimes.
- 它怎么做: Motivated by this, 引入了 CalibAtt, a 训练-free method that accelerates 视频生成 via calibrated sparse attention.
- 得出了什么结果: Extensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled 模型s at various resolutions show that CalibAtt achieves up to 1.58x ...
- 可能的影响: Recent 扩散 模型s enable high-quality 视频生成, but suffer from slow runtimes.
- arXiv: http://arxiv.org/abs/2603.05503v1
3. **Observing and Controlling Features in Vision-Language-Action Models**
- 中文题目(意译): Observing and 控制 特征 in 视觉-语言-动作 模型
- 这篇在讲什么: 视觉-语言-动作 模型s (VLAs) have shown remarkable progress towards embodied intelligence.
- 它怎么做: In this work, 提出了 to close this gap by introducing and analyzing two main concepts: feature-observability and feature-controllability.
- 得出了什么结果: Our 结果显示 that targeted, lightweight interventions can reliably steer a robot's behavior while preserving closed-loop capabilities.
- 可能的影响: 视觉-语言-动作 模型s (VLAs) have shown remarkable progress towards embodied intelligence.
- arXiv: http://arxiv.org/abs/2603.05487v1
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- 中文题目(意译): MM-CondChain: A Programmatically Verified 基准 for Visually Grounded Deep Compositional Reasoning
- 这篇在讲什么: Experiments on a range of MLLMs show that even the strongest 模型 attains only 53.33 Path F1, with sharp drops on hard negatives and as dep...
- 它怎么做: In 本文, 引入了 MM-CondChain, a 基准测试 for visually grounded deep compositional reasoning.
- 得出了什么结果: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- 可能的影响: Multimodal Large Language 模型s (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step de...
- arXiv: http://arxiv.org/abs/2603.12266v1
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- 中文题目(意译): Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- 这篇在讲什么: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- 它怎么做: To address this trade-off, 提出了 Video Streaming Thinking (VST), a novel paradigm for streaming video understanding.
- 得出了什么结果: This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning la...
- 可能的影响: Online Video Large Language 模型s (VideoLLMs) play a critical role in supporting responsive, real-time interaction.
- arXiv: http://arxiv.org/abs/2603.12262v1
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- 中文题目(意译): SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene 生成
- 这篇在讲什么: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- 它怎么做: In 本文, 引入了 SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation.
- 得出了什么结果: At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achi...
- 可能的影响: Text-to-3D scene generation from natural language is highly desirable for digital content creation.
- arXiv: http://arxiv.org/abs/2603.12238v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis**
- arXiv: http://arxiv.org/abs/2603.05483v1
- 类别: cs.LG | HotScore: 23.0 | 作者: Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss
- 速读: SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival An...cs.LG
2. **Accelerating Text-to-Video Generation with Calibrated Sparse Attention**
- arXiv: http://arxiv.org/abs/2603.05503v1
- 类别: cs.CV | HotScore: 18.0 | 作者: Shai Yehezkel, Shahar Yadin, Noam Elata
- 速读: Accelerating Text-to-Video Generation with Calibrated Sparse Attentioncs.CV
3. **Observing and Controlling Features in Vision-Language-Action Models**
- arXiv: http://arxiv.org/abs/2603.05487v1
- 类别: cs.RO | HotScore: 18.0 | 作者: Hugo Buurmeijer, Carmen Amo Alonso, Aiden Swann
- 速读: Observing and Controlling Features in Vision-Language-Action Modelscs.RO
4. **Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation**
- arXiv: http://arxiv.org/abs/2603.05485v1
- 类别: cs.AI | HotScore: 17.0 | 作者: Benjamin Feuer, Lucas Rosenblatt, Oussama Elachqar
- 速读: Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluationcs.AI
5. **An interpretable prototype parts-based neural network for medical tabular data**
- arXiv: http://arxiv.org/abs/2603.05423v1
- 类别: cs.LG | HotScore: 17.0 | 作者: Jacek Karolczak, Jerzy Stefanowski
- 速读: An interpretable prototype parts-based neural network for medical tabular datacs.LG
1. **MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning**
- arXiv: http://arxiv.org/abs/2603.12266v1
- 类别: cs.CV | HotScore: 56.54 | 作者: Haozhan Shen, Shilin Yan, Hongwei Xue
- 速读: MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Composit...cs.CV
2. **Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously**
- arXiv: http://arxiv.org/abs/2603.12262v1
- 类别: cs.CV | HotScore: 56.53 | 作者: Yiran Guan, Liang Yin, Dingkang Liang
- 速读: Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneouslycs.CV
3. **SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation**
- arXiv: http://arxiv.org/abs/2603.12238v1
- 类别: cs.CV | HotScore: 56.46 | 作者: Jun Luo, Jiaxiang Tang, Ruijie Lu
- 速读: SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generationcs.CV
4. **SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning**
- arXiv: http://arxiv.org/abs/2603.12249v1
- 类别: cs.CL | HotScore: 55.5 | 作者: Ziyu Chen, Yilun Zhao, Chengye Wang
- 速读: SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoningcs.CL
5. **Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training**
- arXiv: http://arxiv.org/abs/2603.12246v1
- 类别: cs.AI | HotScore: 53.49 | 作者: Yixin Liu, Yue Yu, DiJia Su
- 速读: Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Trainingcs.AI
## 🆕 最新上新 Top 10
1. Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups (cs.CV) - http://arxiv.org/abs/2603.05507v1
2. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning (cs.CV) - http://arxiv.org/abs/2603.05506v1
3. RoboPocket: Improve Robot Policies Instantly with Your Phone (cs.RO) - http://arxiv.org/abs/2603.05504v1
4. Accelerating Text-to-Video Generation with Calibrated Sparse Attention (cs.CV) - http://arxiv.org/abs/2603.05503v1
5. POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation (cs.LG) - http://arxiv.org/abs/2603.05500v1
6. The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks (cs.AI) - http://arxiv.org/abs/2603.05498v1
7. Safe-SAGE: Social-Semantic Adaptive Guidance for Safe Engagement through Laplace-Modulated Poisson Safety Functions (cs.RO) - http://arxiv.org/abs/2603.05497v1
8. Cheap Thrills: Effective Amortized Optimization Using Inexpensive Labels (cs.LG) - http://arxiv.org/abs/2603.05495v1
9. Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation (cs.LG) - http://arxiv.org/abs/2603.05494v1
10. cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots (cs.RO) - http://arxiv.org/abs/2603.05493v1
1. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation (cs.CV) - http://arxiv.org/abs/2603.12267v1
2. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning (cs.CV) - http://arxiv.org/abs/2603.12266v1
3. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams (cs.CV) - http://arxiv.org/abs/2603.12265v1
4. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing (cs.CV) - http://arxiv.org/abs/2603.12264v1
5. $Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation (cs.RO) - http://arxiv.org/abs/2603.12263v1
6. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously (cs.CV) - http://arxiv.org/abs/2603.12262v1
7. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos (cs.LG) - http://arxiv.org/abs/2603.12261v1
8. HumDex:Humanoid Dexterous Manipulation Made Easy (cs.RO) - http://arxiv.org/abs/2603.12260v1
9. DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning (cs.CV) - http://arxiv.org/abs/2603.12257v1
10. Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (cs.CV) - http://arxiv.org/abs/2603.12255v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。