Skip to main content
buildradar
Sign in
Topic · dataset

dataset

Tracked open-source repos tagged dataset, sorted by stars.

148 repos
  • Public BTC trading context since 2020.

    1,216+1Star change over the last 7 days
  • M2DGR@SJTU-ViSYS

    M2DGR: a Multi-modal and Multi-scenario Dataset for Ground Robots(RA-L2021 & ICRA2022)

    1,201+5Star change over the last 7 days
  • chat-dataset-baseline@hikariming

    人工精调的中文对话数据集和一段chatglm的微调代码

    1,191+0Star change over the last 7 days
  • torchxrayvision@mlmed

    TorchXRayVision: A library of chest X-ray datasets and models. Classifiers, segmentation, and autoencoders.

    1,188+3Star change over the last 7 days
  • covid-19@datasets

    Novel Coronavirus 2019 time series data on cases

    1,165+1Star change over the last 7 days
  • domains@tb0hdan

    World’s single largest Internet domains dataset

    1,157+1Star change over the last 7 days
  • A list of Twitter datasets and related resources.

    1,125+1Star change over the last 7 days
  • SoundMind@xid32

    We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities.

    1,113+0Star change over the last 7 days
  • game-datasets@leomaurodesenv

    :video_game: A curated list of awesome game datasets, and tools to artificial intelligence in games

    1,112+3Star change over the last 7 days
  • insuranceqa-corpus-zh@chatopera

    :helicopter: 保险行业语料库,聊天机器人

    1,064-1Star change over the last 7 days
  • hagrid@hukenovs

    HAnd Gesture Recognition Image Dataset

    1,057+6Star change over the last 7 days
  • TransportationNetworks@bstabler

    Transportation Networks for Research

    1,050+1Star change over the last 7 days
  • MMSA@thuiar

    MMSA is a unified framework for Multimodal Sentiment Analysis.

    1,045+3Star change over the last 7 days
  • name-dataset@philipperemy

    The Python library for names.

    1,019+0Star change over the last 7 days
  • ThoughtSource@OpenBioLink

    A central, open resource for data and tools related to chain-of-thought reasoning in large language models. Developed @ Samwald research group: https://samwald.info/

    1,015+0Star change over the last 7 days
  • A curated, public list of resources for biomechanics and human motion analysis: datasets, processing tools, software for simulation, educational videos, lectures, etc.

    1,010+1Star change over the last 7 days
  • githut@madnight

    Github Language Statistics

    1,004-1Star change over the last 7 days
  • OmniAnomaly@NetManAIOps

    KDD 2019: Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network

    948+5Star change over the last 7 days
  • Mobius@microsoft

    C# and F# language binding and extensions to Apache Spark

    948+0Star change over the last 7 days
  • semantic-segmentation@sithu31296

    SOTA Semantic Segmentation Models in PyTorch

    942-1Star change over the last 7 days
  • tfrecord@vahidk

    Standalone TFRecord reader/writer with PyTorch data loaders

    902+0Star change over the last 7 days
  • SemanticKITTI API for visualizing dataset, processing data, and evaluating results.

    898+3Star change over the last 7 days
  • deepfabric@nolabs-ai

    Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline

    885+3Star change over the last 7 days
  • magpie@magpie-align

    [ICLR 2025] Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing. Your efficient and high-quality synthetic data generation pipeline!

    880+1Star change over the last 7 days
  • OpenCUA@xlang-ai

    [NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents

    833+4Star change over the last 7 days
  • MINT-1T@mlfoundations

    🍃 MINT-1T: A one trillion token multimodal interleaved dataset.

    832+0Star change over the last 7 days
  • PrimeKG@mims-harvard

    Precision Medicine Knowledge Graph (PrimeKG)

    820+3Star change over the last 7 days
  • opendata@NationalGalleryOfArt

    The National Gallery of Art Open Data Program

    818+5Star change over the last 7 days
  • pycococreator@waspinator

    Helper functions to create COCO datasets

    784+0Star change over the last 7 days
  • Total-Text-Dataset@cs-chan

    Total Text Dataset. It consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved, one of a kind.

    772+0Star change over the last 7 days
← Back to topics