Apply now »

Data Science Architect | 9 to 14 years | Gurgaon

At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.

Job Description

Owns the data science lifecycle end-to-end across multi-industry engagements - presales and opportunity shaping, greenfield ML/AI platform architecture, model development and productionization, and program governance through steady-state delivery, with a roadmap toward autonomous, self-optimizing AI/ML operations.

Key Responsibilities

•            Lead presales engagements - RFP/RFI response, AI/ML solution architecture, effort estimation and commercial shaping - and present win themes to executive level stakeholders.

•            Own the end-to-end data science lifecycle - problem framing, exploratory data analysis, feature engineering, model development, validation and deployment.

•            Architect greenfield data and ML platforms (cloud-native data lakes/lakehouses, feature stores, MLOps pipelines), including vendor and tooling selection.

•            Define the roadmap toward autonomous, self-optimizing ML operations - automated retraining, drift detection and closed-loop model monitoring.

•            Lead large-scale brownfield data and analytics platform modernization and legacy-to-target migrations with minimal business disruption.

•            Partner with data engineering to ensure pipeline quality, governance and lineage, while personally owning model design, validation and business impact.

•            Own delivery governance for multi-year, multi-workstream AI/ML transformation programs - scope, schedule, risk, quality and financials.

•            Act as single technical point of accountability across design, build, deploy and operate phases, coordinating data engineering, MLOps and business teams.

•            Establish and track model and program KPIs (accuracy, drift, business impact), steering committee reporting and executive dashboards.

•            Mentor data science/delivery leads and build reusable accelerators, ML frameworks and playbooks across engagements.

Technical Skills & Tools

Category

Tools / Technologies

Languages & ML/DL

Python, R, SQL, Scikit-learn, TensorFlow, PyTorch, XGBoost

MLOps & Deployment

MLflow, Kubeflow, Amazon SageMaker, Azure ML, Vertex AI

Data Platforms

Databricks, Snowflake, Apache Spark, Hadoop

Cloud

AWS, Azure, GCP

Visualization

Power BI, Tableau

Generative AI

LangChain, Azure OpenAI/OpenAI, RAG frameworks

Frameworks & Program Delivery

CRISP-DM, Agile/SAFe, TOGAF, MS Project, Jira, Confluence, PMP/Prince2

 

Primary Skills

Data Science & ML # Experience in designing and deploying advanced analytics and AI solutions using traditional Machine Learning techniques including Classification, Regression, Clustering, Recommendation Systems, Anomaly Detection, Time Series Forecasting, and Reinforcement Learning.S

tatistics # Deep understanding of statistical concepts such as Probability Theory, Hypothesis Testing, Confidence Intervals, Bayesian Statistics, A/B Testing, Experimental Design, Correlation Analysis, Multivariate Statistics, Sampling Techniques, and Predictive

Data Modelling # Modeling. Ability to evaluate data quality, identify bias and fairness concerns, perform causal inference, and develop explainable AI solutions using industry-standard methodologies.

Big Data # Experienced in working with large-scale datasets and Big Data technologies including Spark, Hadoop, Databricks, Kafka, and distributed computing frameworks. Proficient in Python, SQL, and modern data science ecosystems.Deep Learning & GenAI# Deep Learning, Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Prompt Engineering, Vector Databases, Agentic AI frameworks, and MLOps practices for enterprise-scale AI deployments. Demonstrated ability to translate complex business problems into data-driven solutions while ensuring Responsible AI, model governance, transparency, and measurable business outcomes.

Capgemini is a global business and technology transformation partner, helping organizations to accelerate their dual transition to a digital and sustainable world, while creating tangible impact for enterprises and society. It is a responsible and diverse group of 340,000 team members in more than 50 countries. With its strong over 55-year heritage, Capgemini is trusted by its clients to unlock the value of technology to address the entire breadth of their business needs. It delivers end-to-end services and solutions leveraging strengths from strategy and design to engineering, all fueled by its market leading capabilities in AI, generative AI, cloud and data, combined with its deep industry expertise and partner ecosystem.

Ref. code:  549550
Posted on:  5 Oct 2026
Experience Level:  Experienced Professionals
Contract Type:  Permanent
Location: 

Gurgaon, IN

Brand:  Capgemini Engineering
Professional Community:  Manufacturing & Operations Engineering

Apply now »