A DBMS project on Textile Store Management using StreamLit-Python for the frontend app
-
Updated
Feb 26, 2024 - Python
A DBMS project on Textile Store Management using StreamLit-Python for the frontend app
A Python library for intelligent file filtering using SQL expressions and metadata-based scan planning. This library enables efficient data lake query optimization by determining which files need to be scanned based on their statistical metadata.
This repo consists of all the assignments, projects, tasks of Information Retrieval course of FAST NUCES Spring 2023.
Efficient web information retrieval and summation without excessive token usage.
🪂 Parachute: Single-Pass Bi-Directional Information Passing (VLDB'25)
A new package that analyzes user-submitted text queries about community moderation or platform policies (like why small voting projects get flagged) and returns structured insights. It uses an LLM to
Simple search system that includes inverted index builder and boolean query processor for information retrieval.
Data Processing At Scale
Developed an SQL engine that will run a subset of SQL queries using command-line interface
Inverted index and Positional index for a set of collection to facilitate Boolean Model of IR. Inverted files and Positional files are the primary data structure to support the efficient determination of which documents contain specified terms and at which proximity.
Relational database engine built from scratch in Python with SQL-style parsing, hash indexing, aggregations, and multiple join algorithms.
Text preprocessing, indexer constructions, and search engines implementation for information retrieval. Performance analysis done by measuring the construction time of indexers.
A search engine that ranks documents by relevance to a query using a weighting scheme, tokenization, stop word removal, and stemming
A Zarr v3 codec that makes ZFP-compressed arrays in cloud storage queryable without full downloads.
A modular Information Retrieval and Web Mining search engine built in Python, supporting Boolean, phrase, wildcard, proximity, phonetic search, spell correction, and query expansion.
Database Systems
A practical GUI tool for query deduplication with exact dedup, SimHash near-text dedup, and local TF-IDF similarity grouping.
what if I had to make a datalog in a cabin with no internet
How to evaluate Query Optimization based on the IMDb dataset using PostgreSQL and the Join Order Benchmark (JOB).
To associate your repository with the query-processing topic, visit your repo's landing page and select "manage topics."