Skip to main content
buildradar
Sign in

CatchTheTornado/text-extract-api

@CatchTheTornado

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

Stars
3,177
Forks
279
Language
Python
License
MIT
Last push
9 months ago
Pythonjsonpdfllmapipiiextractanonymizationocrocr-python

No related intel yet

This repo has not appeared in any of the sources the radar tracks. The collector runs on a schedule — check back once it covers this repo.