I am a Data Analyst who turns raw, messy data into actionable insights and machine learning models. I work across the full cycle: cleaning and modeling data in SQL and Python, building and explaining predictive models, deploying them with Docker, and delivering results in dashboards. I also build AI agents that answer business questions in SQL.
- Data Analysis & Processing: Python, Pandas, NumPy, SQL, PostgreSQL
- Machine Learning: scikit-learn, XGBoost, K-means, SHAP, MLflow
- Data Visualization: Matplotlib, Seaborn, Plotly
- Dashboards: Power BI (DAX, Power Query), Tableau, Streamlit
- Deployment & Engineering: Docker, FastAPI, pytest, Git
- AI Development: LangChain, LangGraph, AI Agents, Claude Code
- Automation & Extraction: Web Scraping (BeautifulSoup, Requests), Scripting (openpyxl)
🤖 Data Science & Machine Learning
- Customer Churn Prediction & Segmentation Who are an online retailer's customers, and who is about to stop buying? Segmented 5,852 customers with RFM + K-means and predicted churn with XGBoost (ROC-AUC 0.76 on unseen customers), explained with SHAP. Includes MLflow tracking, a FastAPI model served with Docker, and a 4-page Power BI dashboard on PostgreSQL.
🧠 AI Development
- Text-to-SQL Agent AI Agent that converts natural language to SQL queries using LangChain, Google AI Studio (Gemini), and FastAPI.
📊 Data Analysis
- Mexico City Airbnb Analysis End-to-end analysis of 29,329 Airbnb listings in Mexico City: data cleaning in Python, business questions answered in SQL, and results delivered as a Power BI report and an interactive Streamlit app.
- IBM Data Analyst Professional Certificate — Verify Credential
- LinkedIn: José Fernando Avila Camberos