5.6 KiB
5.6 KiB
ArXiv Daily Brief - 2026-03-11
🧠 今日 Top 3(中文可读版)
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成
- 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
- 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin...
- 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41...
- 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
- arXiv: http://arxiv.org/abs/2603.08652v1
- Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
- 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting
- 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence.
- 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ...
- 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
- 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
- arXiv: http://arxiv.org/abs/2603.08707v1
- OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
- 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning
- 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o...
- 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
- arXiv: http://arxiv.org/abs/2603.08655v1
🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- arXiv: http://arxiv.org/abs/2603.08652v1
- 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang
- 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation(cs.AI)
- Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
- arXiv: http://arxiv.org/abs/2603.08707v1
- 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith
- 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting(cs.LG)
- OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
- arXiv: http://arxiv.org/abs/2603.08655v1
- 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins
- 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning(cs.AI)
- Agentic Critical Training
- arXiv: http://arxiv.org/abs/2603.08706v1
- 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho
- 速读: Agentic Critical Training(cs.AI)
- Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines
- arXiv: http://arxiv.org/abs/2603.08704v1
- 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga
- 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...(cs.AI)
🆕 最新上新 Top 10
- Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1
- FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1
- Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1
- Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1
- Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1
- HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1
- A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1
- Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1
- Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1
- Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1
Val 今日建议
- 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
- 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。