Pallavi Bhimte
- Data Scientist
- AI Engineer
- ML Engineer
- Data Analyst
- Data Scientist
Three years in consulting taught me to build data and AI systems that quietly do the hard work, turning a messy problem into something people actually rely on.
Selected work
Client projects at Deloitte and Emmi.
Scoped, built, handed over.
Workers' Compensation Chatbot
An end-to-end AWS Lex conversational system covering 90+ query types across 95 flows. I defined a 12-scenario evaluation framework and led six SMEs through model training, then demoed it live to 13 senior executives. Adopted organisation-wide after the pilot.
Cloud Data Platform
Architected a four-pipeline AWS + dbt platform and migrated 500+ SAP extractors onto it. Worked across nine client stakeholders in product, BI and business, then handed the whole thing to a five-person implementation team with the documentation to run it.
SAS Migration Assessment
Assessed 120+ SAS scripts across nine systems and modelled three migration strategies against value, effort and risk. Presented a three-phase roadmap with four options to senior stakeholders, sized at 5–7 person-months.
Portfolio Carbon Analytics
Built a Python application that quantifies and explains the carbon emissions inside a customer's investment portfolio, turning dense financial data into something a product team could actually reason about. Designed the Snowflake tables underneath it for KPI tracking and reporting.
The path
Journey through 2 continents, 3 companies,
5 industries, and a lot of data.
Things I built on my own
Personal projects, open-source repos.
ResearchMind
Four specialised agents that research a topic together: a Search agent hits Tavily, a Reader agent scrapes the good sources, a Writer chain drafts the report, and a Critic chain scores it and lists what's weak. The critic is the interesting part: it's honest enough to give its own pipeline a 6/10.
View repoAdvanced RAG
A worked series on retrieval, built up from scratch: simple RAG, then embedding models, then semantic chunking, then contextual retrieval. Includes t-SNE projections of the token embeddings and input/output similarity heatmaps, because the failure modes are much easier to see than to describe.
View repoFastAPI + Model Serving
Two APIs. One serves a Random Forest car-price model, loaded once at startup, one-hot encoding handled server-side, with a Streamlit front end to poke at it. The other is an e-commerce REST API built to push Pydantic hard: nested models, computed fields, custom validators, allowlisted seller domains.
View repoStack
Tools I've worked with, over the years.
Languages
- Python
- SQL
- PySpark
- R
GenAI
- LangChain / LangGraph
- Azure OpenAI
- Pinecone / RAG
- FastAPI
- LangSmith / CrewAI
Cloud & data
- AWS
- Databricks
- Snowflake
Tools
- Airflow
- Terraform
- dbt
- GitHub
ML & analytics
- scikit-learn
- Pandas
- NumPy
- A/B & hypothesis testing
Credentials
Check out my verified badges!
Experience
-
Emmi
Dec 2024 to Feb 2025 MelbourneData & ML Engineer
- Built an analytics-driven Python application quantifying and explaining customer carbon emissions from investment portfolios, translating complex financial data into interpretable insights for product and business teams.
- Designed analytics-ready Snowflake tables supporting KPI tracking, dashboards and product reporting.
-
Deloitte
May 2022 to Nov 2024 MelbourneAI & Data Consultant
- Architected an AWS + dbt cloud data platform ingesting 450K+ records/day across 42 sources; migrated 500+ SAP extractors and cut manual reporting effort 70%.
- Built an AWS Lex conversational AI system covering 90+ query types at ~90% production accuracy, reducing routine support query volume 70%; adopted organisation-wide after pilot.
- Assessed 120+ SAS scripts across nine systems, projecting 30% cost and 60% complexity reduction in a three-phase migration roadmap.
- Built PySpark workflows processing 100K+ member records across three release cycles, achieving 98%+ data accuracy and 70% less manual reconciliation.
- Shipped an LLM-powered internal automation app (OpenAI API) standardising seven job task categories across two teams.
Education
-
Master of Data ScienceRMIT University, Melbourne, Mar 2020 to Nov 2021 -
B.Tech, Computer Science & EngineeringGalgotias University, India, Aug 2015 to Dec 2019
About
I turn ambiguous problems into systems people actually use, and it starts with the data. What companies actually pay for isn't the cleverest model, it's one that ships, holds up unattended, and gets explained in language stakeholders can act on. With changing clients, teams and technologies, I've learned what actually matters: understand the real problem before reaching for a solution, then ship something the team can trust.
- Currently building
- Advanced RAG, a from-scratch retrieval pipeline, one failure mode at a time
- Currently learning
- Why retrieval breaks: semantic chunking, contextual retrieval, and where the embeddings actually lie
Off the clock
Beyond work, you'll find me
-
Motorbiking
-
Reading
-
Exploring
-
Building Projects
Let's talk about your messy data.
Open to Data Scientist, AI Engineer, ML Engineer and Data Analyst roles in Melbourne or remote. Australian permanent resident with full working rights, no sponsorship needed.