nlp
Tracked open-source repos tagged nlp, sorted by stars.
- #181
The most accurate natural language detection library for Python, suitable for short text and mixed-language text
★ 1,788+0Star change over the last 7 days - #182★ 1,778+0Star change over the last 7 days
- #183★ 1,774+0Star change over the last 7 days
- #184
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
★ 1,772+0Star change over the last 7 days - #185
自然语言处理、知识图谱、对话系统,大模型等技术研究与应用。
★ 1,766+0Star change over the last 7 days - #186★ 1,763+0Star change over the last 7 days
- #187
Bringing BERT into modernity via both architecture changes and scaling
★ 1,716+0Star change over the last 7 days - #188
Awesome Artificial Intelligence, Machine Learning and Deep Learning as we learn it. Study notes and a curated list of awesome resources of such topics.
★ 1,713+0Star change over the last 7 days - #189
📚 数千篇 AI、LLM、NLP、CV 顶会论文解读,每篇 5 分钟读懂核心思想。
★ 1,709+0Star change over the last 7 days - #190
Graph4nlp is the library for the easy use of Graph Neural Networks for NLP. Welcome to visit our DLG4NLP website (https://dlg4nlp.github.io/index.html) for various learning resources!
★ 1,690+0Star change over the last 7 days - #191★ 1,680+0Star change over the last 7 days
- #192★ 1,672+0Star change over the last 7 days
- #193
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
★ 1,667+0Star change over the last 7 days - #194
Pre-Trained Chinese XLNet(中文XLNet预训练模型)
★ 1,646+0Star change over the last 7 days - #195
:us: a python library for parsing unstructured United States address strings into address components
★ 1,637+0Star change over the last 7 days - #196
A coding-free framework built on PyTorch for reproducible deep learning studies. PyTorch Ecosystem. 🏆26 knowledge distillation methods presented at TPAMI, CVPR, ICLR, ECCV, NeurIPS, ICCV, AAAI, etc are implemented so far. 🎁 Trained models, training logs and configurations are available for ensuring the reproducibiliy and benchmark.
★ 1,630+0Star change over the last 7 days - #197
WikiChat is an improved RAG. It stops the hallucination of large language models by retrieving data from a corpus.
★ 1,615+0Star change over the last 7 days - #198★ 1,602+0Star change over the last 7 days
- #199
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
★ 1,596+0Star change over the last 7 days - #200
Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
★ 1,592+0Star change over the last 7 days - #201
similarity: Text similarity calculation Toolkit for Java. 文本相似度计算工具包,java编写,可用于文本相似度计算、情感分析等任务,开箱即用。
★ 1,584+0Star change over the last 7 days - #202
This is a continuously updated handbook for readers to easily track the latest Text-to-SQL techniques in the literature and provide practical guidance for researchers and practitioners.
★ 1,580+0Star change over the last 7 days - #203
What's in your data? Extract schema, statistics and entities from datasets
★ 1,579+0Star change over the last 7 days - #204
Must-read Papers on Textual Adversarial Attack and Defense
★ 1,575+0Star change over the last 7 days - #205
A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.
★ 1,574+0Star change over the last 7 days - #206
Collection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
★ 1,555+0Star change over the last 7 days - #207★ 1,553+0Star change over the last 7 days
- #208
a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.
★ 1,550+0Star change over the last 7 days - #209★ 1,538+0Star change over the last 7 days
- #210★ 1,507+0Star change over the last 7 days