p
pravashp

Pravash Purkay

@pravashp

Data Engineer

India
Inglés
Parte de la información aparece en idioma inglés.
Sobre mí
I am a Data Scientist with 2+ years of production experience building machine learning models, RAG systems, and LLM-powered applications. I have a proven track record of delivering significant cost savings and efficiency improvements, including a 95% reduction in analytics turnaround time at Pine Labs. I am proficient in Python, SQL, and AWS, with a strong foundation in statistical modelling and NLP.... Lee más

Habilidades

p
pravashp
Pravash Purkay
desconectado • 
Tiempo medio de respuesta: 1 hora

Revisa mis servicios

Desarrollo de chatbots de IA
I will build a custom whatsapp API business agent

Porfolio

Experiencia laboral

Pine_Labs

Data Engineer

Pine Labs • Tiempo completo

Jul 2023 - Jul 20263 yrs

Built and deployed a production-grade NLP system using RAG, LLMs, and semantic search to auto-generate validated SQL queries from natural language inputs across 1,800+ enterprise tables — applying embedding-based retrieval, contextual ranking, and transformer-based text understanding at scale. – Designed end-to-end ML pipelines for enterprise analytics automation using Python, LangChain, LangGraph, and ChromaDB, reducing average query resolution time from 5–7 days to under 30 seconds — a 99%+ latency improvement across 50+ business users. – Applied statistical modelling and anomaly detection techniques to merchant transaction data across 10+ KPIs— identifying revenue trend deviations, refund rate anomalies, and settlement pattern irregularities using SQL window functions and Python-based statistical tests. – Implemented feature engineering and data preprocessing workflows over 56,000+ columns of enterprise transaction data — including missing value imputation, schema normalisation, and embedding-based dimensionality reduction for downstream ML and retrieval tasks. – Built AI-powered self-serve analytics and predictive reporting tools using Streamlit, Plotly, and AWS S3, reducing decision-making turnaround from 1 week to 1 day (60% improvement) and improving ticket resolution efficiency by 95%. – Designed and maintained interactive BI dashboards in Apache Superset tracking merchant performance — used weekly by 5+ cross-functional teams for data-driven strategic decisions. – Led migration of 100+ data tables to Amazon Redshift cloud architecture, optimising query performance by 40% and reducing annual infrastructure costs by approximately $61,560. – Developed model evaluation and validation frameworks for LLM-generated outputs — including benchmark design, accuracy scoring, and failure mode analysis — achieving 87%+ first-attempt accuracy across a 60-question test suite.