# ArXiv Daily Brief - 2026-03-11 ## 🧠 今日 Top 3(中文可读版) 1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation** - 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成 - 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the... - 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin... - 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41... - 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the... - arXiv: http://arxiv.org/abs/2603.08652v1 2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting** - 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting - 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence. - 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ... - 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s. - 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s. - arXiv: http://arxiv.org/abs/2603.08707v1 3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning** - 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning - 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus. - 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus. - 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o... - 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus. - arXiv: http://arxiv.org/abs/2603.08655v1 ## 🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索) 1. **CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation** - arXiv: http://arxiv.org/abs/2603.08652v1 - 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang - 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation(cs.AI) 2. **Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting** - arXiv: http://arxiv.org/abs/2603.08707v1 - 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith - 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting(cs.LG) 3. **OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning** - arXiv: http://arxiv.org/abs/2603.08655v1 - 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins - 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning(cs.AI) 4. **Agentic Critical Training** - arXiv: http://arxiv.org/abs/2603.08706v1 - 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho - 速读: Agentic Critical Training(cs.AI) 5. **Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines** - arXiv: http://arxiv.org/abs/2603.08704v1 - 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga - 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...(cs.AI) ## 🆕 最新上新 Top 10 1. Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1 2. FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1 3. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1 4. Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1 5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1 6. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1 7. A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1 8. Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1 9. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1 10. Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1 ## Val 今日建议 - 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。 - 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。