Skip to main content
buildradar
Sign in
Topic · nlp

nlp

Tracked open-source repos tagged nlp, sorted by stars.

440 repos
  • spacy-stanza@explosion

    💥 Use the latest Stanza (StanfordNLP) research models directly in spaCy

    746+0Star change over the last 7 days
  • Octopii@redhuntlabs

    An AI-powered Personal Identifiable Information (PII) scanner.

    744+0Star change over the last 7 days
  • poetry@sheepzh

    地球上最全的华语现代诗歌语料库,3k+诗人,80K+诗歌,15M+字

    740+2Star change over the last 7 days
  • primeqa@primeqa

    The prime repository for state-of-the-art Multilingual Question Answering research and development.

    739-1Star change over the last 7 days
  • pubmed_parser@titipata

    :clipboard: A Python Parser for PubMed Open-Access XML Subset and MEDLINE XML Dataset

    738+2Star change over the last 7 days
  • Legal-Text-Analytics@Liquid-Legal-Institute

    A list of selected resources, methods, and tools dedicated to Legal Text Analytics.

    735+0Star change over the last 7 days
  • attention_sinks@tomaarsen

    Extend existing LLMs way beyond the original training length with constant memory usage, without retraining

    734-1Star change over the last 7 days
  • PromptKG@zjunlp

    PromptKG Family: a Gallery of Prompt Learning & KG-related research works, toolkits, and paper-list.

    734+0Star change over the last 7 days
  • 729+0Star change over the last 7 days
  • xgen@salesforce

    Salesforce open-source LLMs with 8k sequence length.

    726+0Star change over the last 7 days
  • OpenAI-CLIP@moein-shariatnia

    Simple implementation of OpenAI CLIP model in PyTorch.

    726+0Star change over the last 7 days
  • Revisiting Pre-trained Models for Chinese Natural Language Processing (MacBERT)

    721+2Star change over the last 7 days
  • karukan@togatoga

    Japanese Input Method System for Linux, macOS, Neural Kana-Kanji Conversion Engine

    719+6Star change over the last 7 days
  • searchGPT@michaelthwan

    Grounded search engine (i.e. with source reference) based on LLM / ChatGPT / OpenAI API. It supports web search, file content search etc.

    711+0Star change over the last 7 days
  • vectorflow@dgarnitz

    VectorFlow is a high volume vector embedding pipeline that ingests raw data, transforms it into vectors and writes it to a vector DB of your choice.

    703+0Star change over the last 7 days
  • chat@Decalogue

    基于自然语言理解与机器学习的聊天机器人,支持多用户并发及自定义多轮对话

    702-1Star change over the last 7 days
  • A Full Stack ML (Machine Learning) Roadmap involves learning the necessary skills and technologies to become proficient in all aspects of machine learning, including data collection and preprocessing, model development, deployment, and maintenance.

    701+0Star change over the last 7 days
  • Blackstone@ICLRandD

    :black_circle: A spaCy pipeline and model for NLP on unstructured legal text.

    696+0Star change over the last 7 days
  • WeCron@polyrabbit

    :heavy_check_mark: 微信上的定时提醒 - Cron on WeChat

    691+0Star change over the last 7 days
  • langchain_dart@davidmigloz

    Build LLM-powered Dart/Flutter applications.

    686+0Star change over the last 7 days
  • stanford-openie-python@philipperemy

    Stanford Open Information Extraction made simple!

    682+0Star change over the last 7 days
  • Rankify@DataScienceUIBK

    🔥 Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation 🔥. Our toolkit integrates 40 pre-retrieved benchmark datasets and supports 7+ retrieval techniques, 24+ state-of-the-art Reranking models, and multiple RAG methods.

    682-1Star change over the last 7 days
  • medspacy@medspacy

    Library for clinical NLP with spaCy.

    674+0Star change over the last 7 days
  • Chinese_models_for_SpaCy@howl-anderson

    SpaCy 中文模型 | Models for SpaCy that support Chinese

    673+0Star change over the last 7 days
  • ekphrasis@cbaziotis

    Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

    673+0Star change over the last 7 days
  • datefinder@akoumjian

    Find dates inside text using Python and get back datetime objects

    665+2Star change over the last 7 days
  • semchunk@isaacus-dev

    A fast, lightweight and easy-to-use Python library for splitting text into semantically meaningful chunks.

    664+0Star change over the last 7 days
  • JamSpell@bakwc

    Modern spell checking library - accurate, fast, multi-language

    663+0Star change over the last 7 days
  • pysentimiento@pysentimiento

    A Python multilingual toolkit for Sentiment Analysis and Social NLP tasks

    660+0Star change over the last 7 days
  • obsidian-ava@different-ai

    Quickly format your notes with ChatGPT in Obsidian

    658-2Star change over the last 7 days
← Back to topics