AI 早报 2026-04-16
概览
- 共选入 15 条
- 模型发布:3 条
- 开发工具:3 条
- 产品更新:3 条
- 论文/研究:3 条
- 3DGS / XR 专报:3 条
模型发布
arXiv:2602.11236v2 Announce Type: replace Abstract: Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, o…
arXiv:2604.12978v1 Announce Type: cross Abstract: Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluat…
arXiv:2509.10026v4 Announce Type: replace Abstract: As large vision language models (VLMs) advance, their capabilities in multilingual visual question answerin…
开发工具
arXiv:2604.07656v3 Announce Type: replace-cross Abstract: Hyperspectral imaging (HSI) allows researchers to study plant traits non-destructively. By capturing…
arXiv:2511.22364v2 Announce Type: replace-cross Abstract: Open-vocabulary mobile manipulation (OVMM) requires robots to follow language instructions, navigate,…
arXiv:2604.11839v1 Announce Type: cross Abstract: Autonomous AI agents built on open-source runtimes such as OpenClaw expose every available tool to every sess…
产品更新
arXiv:2604.12257v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, e…
arXiv:2604.12391v1 Announce Type: cross Abstract: In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training accele…
arXiv:2604.12625v1 Announce Type: cross Abstract: High-quality global illumination (GI) in real-time rendering is commonly achieved using precomputed lighting…
论文/研究
arXiv:2604.13030v1 Announce Type: new Abstract: While diffusion models dominate the field of visual generation, they are computationally inefficient, applying…
arXiv:2604.12650v1 Announce Type: new Abstract: Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is active…
arXiv:2604.12565v1 Announce Type: cross Abstract: Robots deployed in unstructured environments must coordinate whole-body motion -- simultaneously moving a mob…
3DGS / XR 专报
Highlight Multiple densification strategies are now supported in gsplat, including - DefaultStrategy(): The original 3DGS densification process. - `Defaul…
Highlight Multi-GPU distributed rasterization is supported! E.g. 4 GPUs could lead to more than 3x speedup as well as 3x less memory usage at the same time.…
AI 早报 2026-04-16
概览
模型发布
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
arXiv:2602.11236v2 Announce Type: replace Abstract: Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, o…
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
arXiv:2604.12978v1 Announce Type: cross Abstract: Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluat…
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
arXiv:2509.10026v4 Announce Type: replace Abstract: As large vision language models (VLMs) advance, their capabilities in multilingual visual question answerin…
开发工具
MVOS_HSI: A Python Library for Preprocessing Agricultural Crop Hyperspectral Data
arXiv:2604.07656v3 Announce Type: replace-cross Abstract: Hyperspectral imaging (HSI) allows researchers to study plant traits non-destructively. By capturing…
BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands
arXiv:2511.22364v2 Announce Type: replace-cross Abstract: Open-vocabulary mobile manipulation (OVMM) requires robots to follow language instructions, navigate,…
Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
arXiv:2604.11839v1 Announce Type: cross Abstract: Autonomous AI agents built on open-source runtimes such as OpenClaw expose every available tool to every sess…
产品更新
Style-Decoupled Adaptive Routing Network for Underwater Image Enhancement
arXiv:2604.12257v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, e…
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
arXiv:2604.12391v1 Announce Type: cross Abstract: In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training accele…
Neural Dynamic GI: Random-Access Neural Compression for Temporal Lightmaps in Dynamic Lighting Environments
arXiv:2604.12625v1 Announce Type: cross Abstract: High-quality global illumination (GI) in real-time rendering is commonly achieved using precomputed lighting…
论文/研究
Generative Refinement Networks for Visual Synthesis
arXiv:2604.13030v1 Announce Type: new Abstract: While diffusion models dominate the field of visual generation, they are computationally inefficient, applying…
Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis
arXiv:2604.12650v1 Announce Type: new Abstract: Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is active…
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
arXiv:2604.12565v1 Announce Type: cross Abstract: Robots deployed in unstructured environments must coordinate whole-body motion -- simultaneously moving a mob…
3DGS / XR 专报
v1.1.0
Highlight Multiple densification strategies are now supported in gsplat, including -
DefaultStrategy(): The original 3DGS densification process. - `Defaul…v1.1.1
What's Changed hot fix Full Changelog: nerfstudio-project/gsplat@v1.1.0...v1.1.1
v1.2.0
Highlight Multi-GPU distributed rasterization is supported! E.g. 4 GPUs could lead to more than 3x speedup as well as 3x less memory usage at the same time.…