Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
-
Updated
Sep 30, 2026 - Python
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
Turn Webpage to LLM friendly input text. Similar to Firecrawl and Jina Reader API. Makes RAG, AI web scraping, image & webpage links extraction easy.
wxpath - declarative web crawling with XPath; a Web Query Language (WQL)
ExtractBench - A Benchmark for Schema-Guided Enterprise Document Extraction
A lightweight MCP (Model Context Protocol) server for integrating ComPDF AI with Claude Desktop, enabling AI-powered intelligent document processing and data extraction from PDFs via natural language.
extract data from html table
Get Lyrics for any songs by just passing in the song name (spelled or misspelled) in less than 2 seconds using this awesome Python Library.
This program extracts insider trading data from the sec website and stores it in excel file for the specified time frame.
Extract audio and other data from the Digitech Trio Plus guitar pedal's SD card
Unofficial Python client for Twitter
Extract structured data from any unstructured web page
A simple UI tool to batch crop images to prepare datasets from images and videos.
Different python utility scripts to help automate mundane/repetitive tasks. Useful for performance testers/data scientist or anyone who wants to automate mundane tasks in python.
A Python module for reading data from a plot provided as SVG file.
Extract data from Octopus mdict (*.mdd, *.mdx) files
A toolkit for extracting elements and visualization for Waymo Open Dataset
This is a library for making batch request to Google Analytics Core Reporting v3 API and extracting data from Google Analytics property into Python 3 data structures.
To associate your repository with the extract-data topic, visit your repo's landing page and select "manage topics."