Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.
-
Updated
Sep 17, 2026 - Java
Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.
Open-source document management platform leveraging AWS managed services. RESTful API for document storage, processing, full-text search, and metadata management. Multi-tenant serverless architecture with auto-scaling... deployed entirely in your AWS account.
A lightweight, framework-agnostic Java library for adding watermarks to various file types, including PDFs and videos
Docling simplifies document processing, parsing diverse formats — including advanced PDF understanding — and providing seamless integrations with the gen AI ecosystem
基于 Spring AI Alibaba 架构的企业级智能客服系统 | RAG + 混合检索(向量+全文)| PostgreSQL pgvector + tsvector | 支持多格式文档处理与意图识别
Open-source RAG backend for document ingestion and AI-powered chat with on-premise LLMs
A RAG (Retrieval-Augmented Generation) backend that lets you ask natural-language questions about uploaded PDF documents and get cited, grounded answers.
Document & OCR text/metadata extraction for DuckDB (Java, Apache Tika)
Multi-language SDKs (TypeScript, Python, Go, Java, C#, Ruby, Rust, Swift, PHP) for AI-powered document processing
AI-powered Product Backlog Generator built with Spring Boot and Spring AI
Microsserviço de assistentes de IA com Spring Boot e Spring AI baseado em RAG. Integra OpenAI e pgvector para ingestão de documentos, busca vetorial e geração de respostas contextualizadas por domínio.
State-machine driven Document Processing Orchestration Service built with Spring Boot. Orchestrates OCR, Document Classification, and Named Entity Recognition pipelines with async processing, retry support, logging, and Dockerized deployment.
Batch digitization tool for handwritten historical documents. Draw a template once — the system crops fields, runs OCR, and applies LLM correction
LambDB, is a lightweight, command-line driven NoSQL database prototype built entirely in Java.
Event-driven file upload & search demo on OpenShift
This project is a document processing tool that converts HWP, PDF, DOCX, and other formats into HTML.
Distributed document-processing platform for asynchronous PDF generation using Spring Boot, RabbitMQ, PostgreSQL, MinIO, Docker, and scalable workers.
Event-driven document processing platform powered by Change Data Capture: PostgreSQL, Debezium, Apache Kafka, Spring Boot and Angular.
AI-powered enterprise knowledge assistant — upload documents, ask questions, and get context-aware answers using RAG, Spring Boot 3.5, Spring AI, and pgvector.
Enterprise Java document processing system with AI extraction, classification and workflow routing
To associate your repository with the document-processing topic, visit your repo's landing page and select "manage topics."