Data ScienceAI/MLAnalytics

Yashraj Singh Rajawat

Data Science AnalystAI/ML Engineer

Turning data into intelligent, automated and business-ready solutions.

3+ years building analytics, scalable data pipelines, machine learning models and AI-powered automation, currently as Assistant Manager – Data Analytics at Niva Bupa Health Insurance.

Portrait of Yashraj Singh Rajawat, Data Science Analyst and AI/ML Engineer, wearing a navy blazer
  • Python
  • SQL
  • Machine Learning
  • AI
  • NLP
  • Databricks
  • BigQuery
  • AWS
  • Tableau
  • Power BI
3+ years
across analytics, ML and AI
1M+
records a day through production ETL
20+ hrs
a week saved with Databricks automation
M.Tech
AI & ML at BITS Pilani, in progress

Professional summary

I'm a data professional with 3+ years of experience across analytics, automation, machine learning and AI. My work sits where business questions meet production data: SQL and Python pipelines, KPI frameworks and insights for sales, finance and operations stakeholders, and models and automation that remove manual effort.

At Niva Bupa Health Insurance I've progressed from Data Analyst Trainee to Assistant Manager – Data Analytics. Along the way I've built ETL pipelines processing 1M+ records a day, Databricks automation saving 20+ hours a week, and an AI-powered benchmarking chatbot that cut manual analysis effort by 60%.

I'm now deepening the machine learning side formally through an M.Tech in Artificial Intelligence & Machine Learning at BITS Pilani.

How the work has evolved

  1. Data analytics

    KPI frameworks, Qlik Sense and BigQuery reporting for sales, finance and operations.

  2. Automation

    SQL/Python ETL at 1M+ records a day and Databricks workflows replacing manual Excel.

  3. Machine learning

    Scikit-Learn models, regression, NLP and time-series projects, and ML-driven portfolio optimisation.

  4. AI solutions

    An AI-powered benchmarking chatbot in production use, and formal study in AI & ML.

Measured impact

Outcomes from production analytics at Niva Bupa and from research-grade ML work.

  • 1M+

    records / day

    processed by SQL and Python ETL pipelines feeding real-time analytics

    Niva Bupa

  • 3 days → 6 hrs

    reporting cycle

    after optimising Qlik Sense dashboards and BigQuery models, about a 92% reduction

    Niva Bupa

  • 60%

    less manual analysis

    from the AI-powered KPI benchmarking chatbot

    Niva Bupa

  • 20+ hrs

    saved / week

    by moving manual Excel processes to Databricks workflows

    Niva Bupa

  • 99.9%

    data accuracy

    through automated data validation checks

    Niva Bupa

  • 1.30 → 1.66

    Sharpe ratio

    as the ML-optimised portfolio universe scaled from 50 to 466 stocks

    Dissertation

Experience

Three roles at Niva Bupa since 2023, from trainee to assistant manager.

Niva Bupa Health Insurance

Noida, IndiaOct 2023 – Present

  1. Data Analyst Trainee

    Oct 2023 – Apr 2024

  2. Senior Executive – Data Analytics

    Apr 2024 – Sep 2025

  3. Assistant Manager – Data Analytics

    Oct 2025 – Present

    Current role

Across these roles

  • Developed an AI-powered benchmarking chatbot with Python, Streamlit and NotebookLM, automating KPI comparison across teams and reducing manual analysis effort by 60%.
  • Built scalable ETL pipelines in SQL and Python integrating multiple sources, processing 1,000,000+ records a day for real-time analytics and automated reporting.
  • Automated workflows on Databricks, saving 20+ hours a week and eliminating manual Excel processes.
  • Optimised Qlik Sense dashboards and BigQuery models, cutting reporting time from 3 days to 6 hours.
  • Designed KPI frameworks and delivered insights to stakeholders across sales, finance and operations.
  • Implemented data validation checks that brought data accuracy to 99.9%.
  • Ran A/B tests and cohort analyses to evaluate campaign performance and inform strategy.
  • Python
  • SQL
  • Databricks
  • BigQuery
  • Qlik Sense
  • Streamlit
  • NotebookLM

AI Variant

Bangalore, IndiaFeb 2023 – Jul 2023

Data Science Trainee

  • Built and tuned machine learning models with Scikit-Learn, improving predictive performance by 8%.
  • Performed EDA, data cleaning, preprocessing and feature engineering to produce reliable, model-ready datasets.
  • Developed Matplotlib and Seaborn visualisations to communicate trends and insights to stakeholders.
  • Python
  • Scikit-Learn
  • Pandas
  • Matplotlib
  • Seaborn

Capabilities and technical stack

Grouped by where each skill has actually been applied.

AI & machine learning

Models that are compared, validated and wired into a decision, not left in a notebook.

  • Regression and classification with Scikit-Learn, including tuning that improved predictive performance by 8%
  • Neural networks and SVR for return forecasting inside a portfolio optimiser
  • NLP pipelines for sentiment classification
  • Time-series forecasting with ARIMA, moving averages and decomposition
  • AI-powered automation: a KPI benchmarking chatbot built with Streamlit and NotebookLM

Analytics & business intelligence

Reliable numbers, delivered fast, to the people who run the business.

  • KPI frameworks and insights for sales, finance and operations
  • Qlik Sense dashboards and BigQuery models that cut reporting from 3 days to 6 hours
  • SQL and Python ETL processing 1M+ records a day
  • Data validation to 99.9% accuracy
  • A/B testing and cohort analysis for campaign strategy

Programming & data

Used across analytics, ETL, ML and automation work.

  • Python
  • SQL
  • Pandas
  • NumPy
  • SciPy
  • Statistics

Machine learning

Regression, classification, feature engineering and model evaluation across projects.

  • Scikit-Learn
  • Regression
  • Classification
  • Feature engineering
  • Model evaluation
  • Neural networks
  • SVR

AI & NLP

Sentiment pipelines and an AI benchmarking assistant in production use.

  • NLP
  • Text preprocessing
  • NotebookLM
  • AI-powered automation

Data science methods

Applied in insurance analytics and research projects.

  • EDA
  • Statistical analysis
  • A/B testing
  • Cohort analysis
  • Time-series analysis
  • Predictive modelling

Data engineering & platforms

Pipelines and workflow automation at 1M+ records a day.

  • Databricks
  • BigQuery
  • ETL
  • Data pipelines
  • Data validation

Cloud

AWS services for storage, query and compute.

  • AWS S3
  • AWS Athena
  • AWS EC2

Business intelligence

Dashboards and reporting for business stakeholders.

  • Qlik Sense
  • Tableau
  • Power BI
  • Excel
  • Matplotlib
  • Seaborn

Tools

Day-to-day build and delivery.

  • Streamlit
  • Jupyter
  • Google Colab
  • GitHub

Featured projects

Production AI work and research-grade machine learning, ordered by relevance.

Built at Niva Bupa Health Insurance

AI-Powered Insurance Benchmarking Chatbot

60% less manual analysis

An AI-powered benchmarking assistant that automates KPI comparison across teams.

Problem
Comparing KPIs across teams depended on manual analysis.
Approach
  • Python back end with a Streamlit interface
  • NotebookLM for AI-assisted analysis of benchmarking material
  • Automated KPI comparison across teams
Outcome
Reduced manual analysis effort by 60%.
  • Python
  • Streamlit
  • NotebookLM
  • KPI analytics
  • Automation

Internal work project; source code is not public.

Dissertation, quantitative research

Portfolio Optimisation: Modern Portfolio Theory vs Machine Learning

1.30 → 1.66 Sharpe ratio

An end-to-end research pipeline comparing Markowitz optimisation with ML-driven portfolios (SVR and neural networks) across S&P 500 universes over a 10-year backtest.

Problem
Test whether ML return forecasts improve portfolio construction over classical mean-variance optimisation, with no ready-made dataset.
Approach
  • Automated data collection from Yahoo Finance, FRED and Wikipedia with retry and rate-limit handling
  • Engineered technical indicators (RSI, MACD, moving averages)
  • Implemented Markowitz mean-variance optimisation from scratch with SciPy
  • Trained per-stock SVR and MLP models whose forecasts feed the same optimiser for a like-for-like comparison
  • Traced a NaN-propagation bug and added coverage screening for stocks with partial price history
Outcome
Sharpe ratio rose from 1.30 to 1.48 to 1.66 as the universe grew from 50 to 100 to 466 stocks. The data-integrity fix corrected a result that had been overstated by roughly 35%.
  • Python
  • pandas
  • NumPy
  • SciPy
  • scikit-learn
  • SVR
  • Neural networks
  • Backtesting

NLP project

Social Media Sentiment Analysis

An end-to-end NLP pipeline that classifies social media sentiment and surfaces trends for marketing strategy.

Problem
Turn unstructured social media text into sentiment signals marketing teams can act on.
Approach
  • Text preprocessing and feature engineering
  • Sentiment classification with machine learning and deep learning models
  • Trend analysis across the classified corpus
Outcome
Identified sentiment trends and delivered actionable insights for marketing strategy optimisation.
  • Python
  • NLP
  • Text preprocessing
  • Machine learning
  • Deep learning

Time-series forecasting

Stock Price Prediction (Tesla)

Time-series forecasting on Tesla price data, comparing classical approaches on error metrics.

Problem
Forecast Tesla share prices and establish which classical method performs best.
Approach
  • Seasonal decomposition of the price series
  • Moving-average baselines
  • ARIMA modelling
Outcome
Evaluated and compared multiple forecasting approaches using MAE and RMSE.
  • Python
  • ARIMA
  • Moving averages
  • Decomposition
  • MAE / RMSE

Regression modelling

Electric Synchronous Motor Speed Prediction

100K+ records modelled

Regression models predicting electric motor speed, trained and evaluated on 100,000+ records.

Problem
Predict motor speed accurately from a large set of correlated input features.
Approach
  • Feature engineering and feature selection
  • Multicollinearity analysis with VIF
  • Decision Tree and Neural Network regressors, compared head to head
Outcome
Improved model performance through feature selection and VIF; models evaluated with RMSE and R².
  • Python
  • Regression
  • Decision Trees
  • Neural networks
  • VIF
  • RMSE / R²

Classification modelling

Bankruptcy Prevention

Classification models that flag bankruptcy risk, from EDA through model comparison to a deployment script.

Problem
Predict which companies are at risk of bankruptcy so preventive action can be taken.
Approach
  • Exploratory data analysis
  • Multiple classification algorithms trained and compared
  • Evaluation with classification metrics
  • Model saved and wrapped in a deployment script
Outcome
A compared set of classifiers with the selected model saved for deployment.
  • Python
  • Classification
  • Model comparison
  • EDA

Machine failure classification

Incident Prediction

A classification model predicting machine failure incidents from operational data.

Problem
Anticipate machine failures before they become incidents.
Approach
  • Normalisation and outlier treatment on maintenance data
  • Class balancing for rare failure events
  • Classification modelling
  • Trained model saved and served through a deployment script
Outcome
A machine-failure classifier trained on cleaned, balanced data, saved and packaged for deployment.
  • Python
  • Classification
  • Data balancing
  • Outlier treatment

GitHub projects

Public repositories, ranked by relevance to data science and machine learning rather than by date. 30 public repositories in total.

View all GitHub repositories

Synced from GitHub on 23 Sep 2026.

Education

2026 – PresentIn progress

BITS Pilani

M.Tech in Artificial Intelligence & Machine Learning

Work Integrated Learning ProgrammeRemote, India

2019 – 2023

Rustamji Institute of Technology

Bachelor of Technology (B.Tech) in Information Technology

CGPA 8.32 / 10Gwalior, India

Professional journey

  • Work
  • Education
  1. 2019 – 2023

    B.Tech, Information Technology

    Rustamji Institute of Technology (education)

  2. Feb 2023

    Data Science Trainee

    AI Variant (employment)

  3. Oct 2023

    Data Analyst Trainee

    Niva Bupa Health Insurance (employment)

  4. Apr 2024

    Senior Executive – Data Analytics

    Niva Bupa Health Insurance (employment)

  5. Oct 2025

    Assistant Manager – Data Analytics

    Niva Bupa Health Insurance (employment)

  6. 2026

    M.Tech, AI & Machine Learning

    BITS Pilani (education)

Certifications

Data science, machine learning, AI, SQL and Python credentials.

View certifications on LinkedIn

Resume

Download my latest resume for a detailed overview of my experience, technical skills, projects and education.

Let's Build Something Intelligent.

Open to opportunities across Data Science, Data Analytics, AI/ML and intelligent automation.

Connect with Yashraj Singh Rajawat, a Data Science and AI/ML professional in India, based in Noida, India.