Site Reliability Engineer (SRE)
Keeping AI applications reliable, observable, and maintainable.
I support and operate AI applications at Sber, focusing on reliability, observability, and automation.
- Keep production services stable through clear operational practices and useful monitoring.
- Build practical tools in Python and C# when existing systems need glue, guardrails, or a clean operator experience.
- Keep infrastructure explicit with Terraform, Ansible, containers, and repeatable deployment flows.
- Prefer boring reliability, clear ownership, and systems that can be debugged at 03:00 without guesswork.
| Project | What it does | Stack |
|---|---|---|
| Cleanarr | Cascade media cleanup for Radarr, Sonarr, Jellyseerr, qBittorrent, and Jellyfin. | Python, automation |
- Site reliability: dependable operation and maintenance of AI applications.
- Observability: metrics, logs, and tracing that make production behavior easier to understand.
- Automation: CI/CD, GitOps, release flows, and routine operations with fewer manual steps.
- Self-hosted systems: observable, maintainable services that can be debugged quickly.
- Developer tooling: small CLIs, integrations, and internal tools that remove operational friction.

