Skip to content
@redai-studio

RAI Studio

Building the infrastructure for large model training, inference, optimization, and serving — empowering creators and developers to harness AI at scale.

🤖 RAI Studio · Xiaohongshu AI Infrastructure

Building Leading & Developer-Friendly AI Large Model Full-Stack Infrastructure

Open Source Community AI Infra

🏠 Who We Are

We are the AI Infrastructure Team at Xiaohongshu (Rednote), responsible for the company's end-to-end infrastructure for large models.

We build across three layers — compute, frameworks, and platforms — to support model training, compression, deployment, online and offline serving, and agent development and publishing. Our work helps teams turn model capabilities into production applications with greater efficiency, lower cost, reliable operation, and repeatable delivery at scale.

🚀 Our Mission

Build Xiaohongshu's foundation for productivity in the AI era — making AI as reliable, efficient, and accessible as water and electricity for every business scenario.

We believe better models need infrastructure that consistently turns their capabilities into real-world value. We focus on:

  • 💡 Faster experimentation and delivery — Help teams validate ideas, produce models, and deploy AI applications through a unified toolchain.
  • 🪶 Efficient compute and execution — Improve resource utilization through unified GPU scheduling, heterogeneous hardware support, and training and inference optimization.
  • 🧮 Reliable services at scale — Build dependable model services and observable agent applications that can grow with demand.
  • 🌍 Open collaboration and research — Share reusable systems and research with the community, and advance AI infrastructure together.

📦 Our Infrastructure Stack

RAI Studio connects the compute foundation, model frameworks, and developer platforms through five complementary layers:

Layer Products & Capabilities What It Enables
Agent applications AgentSphere platform: agent orchestration, workflows and tools, knowledge and memory, application publishing, observability and evaluation Build, publish, and operate agent applications
Model services Red Token Hub: online and offline model services, model onboarding, routing and scheduling, KV cache reuse, high availability, and elastic scaling Reliable and efficient Model-as-a-Service (MaaS)
Model production QuickSilver: data management, training, compression, deployment, and evaluation A unified toolchain across the model lifecycle
Frameworks & runtimes 🚄 Relax, ✂️ RedSlim, 🚀 rLLM: our framework portfolio spanning model training, compression, and inference Accelerate model production and execution across heterogeneous hardware
Compute infrastructure Unified management and scheduling across accelerator types and regions, elastic resource allocation, and cluster efficiency optimization Scalable compute capacity and better resource utilization

📦 Research & Open Source

Our open-source work spans reinforcement learning, model compression, efficient inference, reasoning, and computer-use agents.

Project Focus Links
Relax Asynchronous reinforcement learning for omni-modal post-training at scale Code · Paper
HiSVD Hierarchical low-rank model compression guided by information capacity and spectral structure Code · Paper
PIPO — Pair-In, Pair-Out Latent multi-token prediction for efficient large language models Code · Paper
Hint Tuning Improving reasoning with less training data Code · Paper
HYPIC Position-independent caching for faster hybrid-attention model serving Code · Paper
Hybrid Routing Agent Tool use and multimodal context management for hybrid GUI–MCP computer-use agents Code · Paper

🤝 Get Involved

We welcome developers and researchers working on efficient, reliable AI infrastructure.

  • 🚀 Try a project: Start with its README for setup, examples, and supported configurations.

  • 🐛 Report a problem: Open an issue in the relevant repository with reproduction steps.

  • 💬 Share an idea: Open an issue with a concrete proposal or feedback.

  • 🔧 Contribute: Submit code, documentation, examples, or reproducible benchmarks, following the repository's contribution guidance where available.

  • 📚 Build on our research: Read the linked papers and use each project's citation instructions when referencing the work.

  • 🌟 Spread the word: Star and share projects you find useful.

Browse all RAI Studio repositories to find a project that matches your interests.


⭐ Star our projects if you find them helpful!

Made with ❤️ by the Xiaohongshu AI Infrastructure Team

Pinned Loading

  1. Relax Relax Public

    An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

    Python 616 150

  2. PIPO PIPO Public

    Implementation of an efficient LLM architecture: the Pair-In / Pair-Out Model (PIPO)

    Python 43

  3. HiSVD HiSVD Public

    [ACL 2026] HiSVD: Principled Low-Rank Approximation of LLMs via Hierarchical Modeling of Information Capacity and Spectral Structure

    Python 2

  4. HYPIC HYPIC Public

    Position-independent caching for hybrid-attention LLM serving.

    Python 20 6

  5. hint-tuning hint-tuning Public

    Official code, data, and models for "Hint Tuning: Less Data Makes Better Reasoners"

    Python 22 2

  6. hybrid-routing-agent hybrid-routing-agent Public

    Python 17 1

Repositories

Showing 10 of 17 repositories
  • Relax Public

    An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

    redai-studio/Relax's past year of commit activity
    Python 616 Apache-2.0 150 91 72 Updated Sep 17, 2026
  • community Public

    RedAI-Infra Developer Community

    redai-studio/community's past year of commit activity
    Shell 5 Apache-2.0 3 0 0 Updated Sep 17, 2026
  • .github Public

    Xiaohongshu AI Platform Team

    redai-studio/.github's past year of commit activity
    0 0 0 0 Updated Sep 16, 2026
  • TransferQueue Public Forked from Ascend/TransferQueue

    This is the official **live** mirror of https://gitcode.com/Ascend/TransferQueue. Feel free to contribute!

    redai-studio/TransferQueue's past year of commit activity
    Python 0 Apache-2.0 51 0 1 Updated Sep 10, 2026
  • comminsight Public

    CommInsight diagnoses hangs, stragglers, and failures in large-scale distributed AI workloads through multidimensional training semantics reconstruction from low-level communication events and cross-rank analysis.

    redai-studio/comminsight's past year of commit activity
    Go 9 Apache-2.0 0 0 0 Updated Sep 9, 2026
  • redai-studio/hybrid-routing-agent's past year of commit activity
    Python 17 Apache-2.0 1 0 0 Updated Aug 5, 2026
  • HYPIC Public

    Position-independent caching for hybrid-attention LLM serving.

    redai-studio/HYPIC's past year of commit activity
    Python 20 Apache-2.0 6 4 4 Updated Jul 27, 2026
  • sglang Public Forked from sgl-project/sglang

    SGLang is a high-performance serving framework for large language models and multimodal models.

    redai-studio/sglang's past year of commit activity
    Python 0 Apache-2.0 8,978 0 0 Updated Jul 15, 2026
  • hint-tuning Public

    Official code, data, and models for "Hint Tuning: Less Data Makes Better Reasoners"

    redai-studio/hint-tuning's past year of commit activity
    Python 22 2 0 0 Updated Jul 8, 2026
  • PIPO Public

    Implementation of an efficient LLM architecture: the Pair-In / Pair-Out Model (PIPO)

    redai-studio/PIPO's past year of commit activity
    Python 43 Apache-2.0 0 0 0 Updated Jun 10, 2026

Most used topics

Loading…