microsoft/LLMLingua
@microsoft[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
Stars
6,623
Forks
420
Language
Python
License
MIT
Last push
5 months ago
Related intel (0)
No related intel yet
This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.