chore: allowlist trusted plugins
This commit is contained in:
@@ -0,0 +1,62 @@
|
||||
# ArXiv Daily Brief - 2026-03-11
|
||||
|
||||
## 🧠 今日 Top 3(中文可读版)
|
||||
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
|
||||
- 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成
|
||||
- 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
|
||||
- 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin...
|
||||
- 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41...
|
||||
- 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
|
||||
- arXiv: http://arxiv.org/abs/2603.08652v1
|
||||
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
|
||||
- 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting
|
||||
- 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence.
|
||||
- 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ...
|
||||
- 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
|
||||
- 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
|
||||
- arXiv: http://arxiv.org/abs/2603.08707v1
|
||||
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
|
||||
- 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning
|
||||
- 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
|
||||
- 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
|
||||
- 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o...
|
||||
- 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
|
||||
- arXiv: http://arxiv.org/abs/2603.08655v1
|
||||
|
||||
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
|
||||
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
|
||||
- arXiv: http://arxiv.org/abs/2603.08652v1
|
||||
- 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang
|
||||
- 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation(cs.AI)
|
||||
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
|
||||
- arXiv: http://arxiv.org/abs/2603.08707v1
|
||||
- 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith
|
||||
- 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting(cs.LG)
|
||||
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
|
||||
- arXiv: http://arxiv.org/abs/2603.08655v1
|
||||
- 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins
|
||||
- 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning(cs.AI)
|
||||
4. **Agentic Critical Training**
|
||||
- arXiv: http://arxiv.org/abs/2603.08706v1
|
||||
- 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho
|
||||
- 速读: Agentic Critical Training(cs.AI)
|
||||
5. **Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines**
|
||||
- arXiv: http://arxiv.org/abs/2603.08704v1
|
||||
- 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga
|
||||
- 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...(cs.AI)
|
||||
|
||||
## 🆕 最新上新 Top 10
|
||||
1. Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1
|
||||
2. FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1
|
||||
3. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1
|
||||
4. Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1
|
||||
5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1
|
||||
6. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1
|
||||
7. A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1
|
||||
8. Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1
|
||||
9. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1
|
||||
10. Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1
|
||||
|
||||
## Val 今日建议
|
||||
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
|
||||
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。
|
||||
Reference in New Issue
Block a user