Files
val-blog/org/cases/arxiv_digest/output/2026-03-11.md
T

5.6 KiB
Raw Blame History

ArXiv Daily Brief - 2026-03-11

🧠 今日 Top 3(中文可读版)

  1. CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
    • 中文题目(意译): CoCo: Code as CoT for Text-to-Image Preview and Rare Concept 生成
    • 这篇在讲什么: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
    • 它怎么做: In this work, 提出了 CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enablin...
    • 得出了什么结果: Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41...
    • 可能的影响: Recent advancements in Unified Multimodal 模型s (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the...
    • arXiv: http://arxiv.org/abs/2603.08652v1
  2. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
    • 中文题目(意译): Impermanent: A Live 基准 for Temporal Generalization in Time Series Forecasting
    • 这篇在讲什么: While these 模型s often claim broad generalization, existing evaluation protocols provide limited evidence.
    • 它怎么做: 引入了 Impermanent, a live 基准测试 that evaluates forecasting 模型s under open-world temporal change by scoring forecasts sequentially over time ...
    • 得出了什么结果: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
    • 可能的影响: Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style 模型s.
    • arXiv: http://arxiv.org/abs/2603.08707v1
  3. OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
    • 中文题目(意译): OfficeQA Pro: An Enterprise 基准 for End-to-End Grounded Reasoning
    • 这篇在讲什么: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
    • 它怎么做: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
    • 得出了什么结果: Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying o...
    • 可能的影响: 引入了 OfficeQA Pro, a 基准测试 for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus.
    • arXiv: http://arxiv.org/abs/2603.08655v1

🔥 今日热度 Top 5(新鲜度+关键词+HN提及+代码线索)

  1. CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
    • arXiv: http://arxiv.org/abs/2603.08652v1
    • 类别: cs.AI | HotScore: 63.47 | 作者: Haodong Li, Chunmei Qing, Huanyu Zhang
    • 速读: CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generationcs.AI
  2. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
    • arXiv: http://arxiv.org/abs/2603.08707v1
    • 类别: cs.LG | HotScore: 57.86 | 作者: Azul Garza, Renée Rosillo, Rodrigo Mendoza-Smith
    • 速读: Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecastingcs.LG
  3. OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
    • arXiv: http://arxiv.org/abs/2603.08655v1
    • 类别: cs.AI | HotScore: 55.52 | 作者: Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins
    • 速读: OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoningcs.AI
  4. Agentic Critical Training
    • arXiv: http://arxiv.org/abs/2603.08706v1
    • 类别: cs.AI | HotScore: 51.85 | 作者: Weize Liu, Minghui Liu, Sy-Tuyen Ho
    • 速读: Agentic Critical Trainingcs.AI
  5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines
    • arXiv: http://arxiv.org/abs/2603.08704v1
    • 类别: cs.AI | HotScore: 51.85 | 作者: Akshay Gulati, Kanha Singhania, Tushar Banga
    • 速读: Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting...cs.AI

🆕 最新上新 Top 10

  1. Scale Space Diffusion (cs.CV) - http://arxiv.org/abs/2603.08709v1
  2. FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models (cs.CV) - http://arxiv.org/abs/2603.08708v1
  3. Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting (cs.LG) - http://arxiv.org/abs/2603.08707v1
  4. Agentic Critical Training (cs.AI) - http://arxiv.org/abs/2603.08706v1
  5. Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines (cs.AI) - http://arxiv.org/abs/2603.08704v1
  6. HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising (cs.CV) - http://arxiv.org/abs/2603.08703v1
  7. A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies (cs.AI) - http://arxiv.org/abs/2603.08692v1
  8. Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training (cs.LG) - http://arxiv.org/abs/2603.08687v1
  9. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (cs.SD) - http://arxiv.org/abs/2603.08683v1
  10. Structural Causal Bottleneck Models (stat.ML) - http://arxiv.org/abs/2603.08682v1

Val 今日建议

  • 先读 Top 5 里的 1-2 篇,优先看是否有可直接复用的方法/代码。
  • 若你愿意,我下一步可对 Top 3 产出“中文三段式精读卡”(问题-方法-可落地点)。