Torchserve server using a YoloV5 model running on docker with GPU and static batch inference to perform production ready and real time inference.
-
Updated
Feb 10, 2023 - Python
Torchserve server using a YoloV5 model running on docker with GPU and static batch inference to perform production ready and real time inference.
Production-ready batch inference for Cohere’s Arabic/English ASR model, with optimized Silero VAD, multi-file GPU batching, subtitles, and optional word-level timestamps.
Optimize and run consistent AI Functions.
RWKV7 推理服务 | 高并发批量推理,动态扩缩容,兼容OpenAI API,CUDA优化,低显存部署大模型推理 RWKV linear attention batch inference server — auto-queue, dynamic capacity, OpenAI compatible.
Batch inference for datasets: rows in, any OpenAI-compatible endpoint, resumable parquet out. Congestion-aware concurrency — no tuning.
Serve pytorch inference requests using batching with redis for faster performance.
Self-hosted batch LLM pipeline for analyzing customer feedback from Excel. Upload xlsx, describe the task, configure output fields — get structured results. Works with any OpenAI-compatible API and Ollama.
Fake Anthropic / Bedrock (InvokeModel, Converse, batch inference with a bundled S3 emulation) / OpenAI server for zero-cost human/AI-in-the-loop LLM debugging, a scripted test harness (scenario rules, latency and rate-limit simulation, pytest fixture) and a cross-provider relay bridge
Tutorial: mass-parallelized RAI compliance checks with vLLM batch inference on Cloud TPU v5e, using Gemma 4 as an LLM-as-a-Judge
Queue-driven, scale-from-zero GPU inference for any Kubernetes — bursts to cross-region VMs when GPUs run dry
Fine-tune anime-seg (ISNetDIS) bằng LoRA trên 1306 cặp ảnh-mask để tách nền anime; có split cố định, evaluation trên mask thô, hậu xử lý alpha_gain cho RGBA cutout, merge adapter vào checkpoint, và CLI rmbg cho train/eval/predict/report/make-split với manifest tái lập đầy đủ.
AWS-oriented ML platform for fraud detection, churn prediction, anomaly detection, experimentation, monitoring, and risk inference.
End-to-end AWS SageMaker model training, real-time deployment, and batch inference workflow
RankShift Serving — batch engagement prediction with per-row fault isolation
Production-style batch ML pipeline for manufacturing claims anomaly detection with drift checks (synthetic data).
Production-ready ML system for credit default prediction on transactional data (458k clients). Features end-to-end pipeline: 1,158 engineered features, ablation & Top-500 pruning, LightGBM HPO (AMEX 0.791, Gini 0.923), multi-seed stability, CLI batch inference (20.1s/458k), and 267/267 automated tests.
Summarise payment backlogs and render an audit-ready PDF digest through one Infrai credential.
B2B Fintech SaaS Churn Intelligence & Customer Lifetime Value (CLV) Engine with Kimball Star Schema, LightGBM, Tweedie Regressor, and TreeSHAP.
A lightweight, production-style Kubernetes setup that deploys, mounts volumes to, and executes a containerized Python ML inference job on a local cluster using KinD (Kubernetes in Docker).
Offline trip-safety GUI and batch inference with a 19-feature decision-tree pipeline, 63 tests, and a reproducible ~78.8 MB Windows executable smoke test.
To associate your repository with the batch-inference topic, visit your repo's landing page and select "manage topics."