Files
val-blog/org/cases/arxiv_digest/output/2026-03-11.md
T

63 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ArXiv Daily Brief - 2026-03-11
## 🧠 今日 Top 3(中文可读版)
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
- 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成
- 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
- 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin...
- 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41...
- 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
- arXiv: http://arxiv.org/abs/2603.08652v1
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
- 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting
- 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence.
- 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ...
- 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
- 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
- arXiv: http://arxiv.org/abs/2603.08707v1
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
- 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning
- 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o...
- 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- arXiv: http://arxiv.org/abs/2603.08655v1
## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation**
- arXiv: http://arxiv.org/abs/2603.08652v1
- 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang
- 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generationcs.AI
2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting**
- arXiv: http://arxiv.org/abs/2603.08707v1
- 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith
- 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecastingcs.LG
3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning**
- arXiv: http://arxiv.org/abs/2603.08655v1
- 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins
- 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoningcs.AI
4. **Agentic Critical Training**
- arXiv: http://arxiv.org/abs/2603.08706v1
- 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho
- 速读: Agentic Critical Trainingcs.AI
5. **Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines**
- arXiv: http://arxiv.org/abs/2603.08704v1
- 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga
- 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...cs.AI
## 🆕 最新上新 Top 10
1. Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1
2. FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1
3. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1
4. Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1
5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1
6. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1
7. A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1
8. Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1
9. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1
10. Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1
## Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。