Skip to content

version 0.2.1 iteration plan #110

Description

Estimated Release Date: 3/19
Release Manager: Guoxin (@suiguoxin)
Schedule:

  • Design Review: 3/19
  • Coding: 3/10
  • Testing: 3/19

Features

Backlog

  • P1 exp: target comp ratio v.s. real comp ratio on specific data
  • >token level, < sentence level, list different mappings and design interface P1 word level compression When I use chinese llama, the compressed prompt has garbled code #4
  • P1 Support more / faster engines Support for llama.cpp or exl2 #41, including llama_cpp, FasterTransformer, vLLM ETA: TBD
    • survey which engines to support
  • P2 Documentation and examples
    • Supported models and experiment results (with compressor throughput) after a faster engine supported

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions