📄 2026-08-16 arXiv 论文速递
共收录 22 篇与追踪领域相关的新论文。
1. Intern-S2-Preview: Scientific Agentic Foundation Model
- arXiv: 2608.13505v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du 等
- 标签:
cs.LGcs.CLcs.CV - 一句话概述: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models
2. Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
- arXiv: 2608.13239v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera, Stéphane Lathuilière
- 标签:
cs.CV - 一句话概述: Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point. In this context, we
3. DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
- arXiv: 2608.13489v1 · PDF
- 重要度: ⬜ 泛读
- 作者: DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang
- 标签:
cs.CVcs.RO - 一句话概述: We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone
4. ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models
- arXiv: 2608.13438v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Gehan Zheng, Matthew Johnson-Roberson, Weiming Zhi
- 标签:
cs.ROcs.AIcs.CV - 一句话概述: Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-camera setups: close gripper--object views help observe contact, but a poor approach may already push, miss, slip, or disturb the object before conventional detec
5. PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
- arXiv: 2608.13552v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao 等
- 标签:
cs.CV - 一句话概述: Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human p
6. Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
- arXiv: 2608.13341v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang 等
- 标签:
cs.LGcs.AI - 一句话概述: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning met
7. Into the ORBIT for Time Series: Training Regimes for Foundation Models
- arXiv: 2608.13262v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong
- 标签:
cs.LGcs.AI - 一句话概述: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. As a result, pre-training distributions are often poorly controlled with respect to domain imbalance, context requiremen
8. GS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model Priors
- arXiv: 2608.13502v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Yanming Yang, Chenxi Song, Ping Wang, Xin Yuan, Chi Zhang
- 标签:
cs.CV - 一句话概述: Snapshot Compressive Imaging (SCI) offers an efficient solution for high-speed video acquisition and, under exposure-time camera--scene relative motion, multi-view scene capture by compressing temporal or spatial information into a single 2D measurement. While recent studies have explored SCI for 3D
9. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
- arXiv: 2608.13463v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck
- 标签:
cs.CVcs.AIcs.CLcs.LG - 一句话概述: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large lang
10. A Unifying Perspective on Causal World Models: From Observations to Representations to Structure
- arXiv: 2608.13456v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Avinash Kori, Fabrizio Russo
- 标签:
cs.AIcs.CV - 一句话概述: World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptua
11. UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models
- arXiv: 2608.13453v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Yukun Dai, Mingzhe Dai, Tianshi Wang, Fengling Li, Jingjing Li, Lei Zhu
- 标签:
cs.CVcs.AI - 一句话概述: Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them to adversarial interference that may cause unsafe
12. Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
- arXiv: 2608.13441v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ruizhu Zhou, Lingcheng Meng, Lining Hu, Ting Liu, Yuzhuo Fu
- 标签:
cs.CV - 一句话概述: Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code presents a greater challenge: a model must jointly recover visual structure, ground the requested change, generate compilabl
13. Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services
- arXiv: 2608.13315v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu
- 标签:
cs.GTcs.AIcs.LGeess.SY - 一句话概述: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency. We model this int
14. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
- arXiv: 2608.13555v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin 等
- 标签:
cs.ROcs.AIcs.CV - 一句话概述: Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contact
15. A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms
- arXiv: 2608.13511v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Edward Holmberg, Elias Ioup, Mahdi Abdelguerfi
- 标签:
cs.RO - 一句话概述: Repeated in-situ evaluation of ocean-glider planners requires scarce vehicles, operators, deployment and recovery resources, and ocean conditions that cannot be reset for competing algorithms. We present a guided, installation-free browser-native digital test range that transforms a selected region
16. Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning
- arXiv: 2608.13422v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Zheyu Zhuang, Ruiyu Wang, Nick Heppert, Johannes Fabian Hahn, Abhinav Valada, Florian T. Pokorny, Danica Kragic
- 标签:
cs.RO - 一句话概述: Visual bottlenecks that focus policy inputs on regions of interest (ROIs) can improve data-efficient visuomotor learning by separating where to look from how to act. Many ROI interfaces rely on external spatial labels, such as gaze, object classes, or affordance annotations. Label-free alternatives
17. NestDex: Nested Policy Learning with Copilot Assisted Teleoperation for Dexterous Manipulation
- arXiv: 2608.13362v1 · PDF
- 重要度: ⬜ 泛读
- 作者: James Zhao, Jinhe Tang, Mingyuan Ba, Weiming Zhi
- 标签:
cs.RO - 一句话概述: Dexterous manipulation promises substantially richer robot interaction with the physical world, but learning these behaviours remains constrained by the difficulty of collecting consistent, complete-task demonstrations. Unlike parallel-jaw manipulation, dexterous tasks require the operator to coordi
18. Predictive Relative-Velocity Steering for Safe Robotic Manipulator Teleoperation in Dynamic Environments
- arXiv: 2608.13284v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Changhao Hu, Zeyi Liu, Songqiao Hu, Shuang Liu, Zihan Meng, Xiao He
- 标签:
cs.RO - 一句话概述: Recent advances in teleoperation have enabled robotic manipulators to perform dexterous, human-arm-like motions. However, human operators may fail to avoid suddenly appearing obstacles promptly and effectively, particularly under network latency or limited attention, thereby creating safety risks. T
19. V-RAE: Rethinking Video Latent Spaces for Generation
- arXiv: 2608.13556v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Minghui Guo, Shengqiong Wu, Hao Fei
- 标签:
cs.CV - 一句话概述: Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic o
20. Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
- arXiv: 2608.13546v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao
- 标签:
cs.CV - 一句话概述: Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length
21. TabSOM: A tabular-to-image encoding method based on self-organizing maps
- arXiv: 2608.13513v1 · PDF
- 重要度: ⬜ 泛读
- 作者: David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara, Francisco J. Lara-Abelenda, Luis Zhinin-Vera, Diego H. Peluffo-Ordóñez
- 标签:
cs.CVcs.LG - 一句话概述: Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction
22. LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles
- arXiv: 2608.13450v1 · PDF
- 重要度: ⬜ 泛读
- 作者: Md Wasiul Haque, Sagar Dasgupta, Mizanur Rahman, Md Rayhanur Rahman
- 标签:
cs.SEcs.CRcs.LG - 一句话概述: Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming exploitability requires executable test artifacts