AI Data Engineer
Who You Will Be Working With
As a Data Engineer at Capgemini, you'll design, build and operate reliable, scalable data pipelines and data products: the trusted foundations that power analytics, AI, GenAI and increasingly agentic systems. You'll work hands-on with modern data engineering tooling and cloud services to ingest, transform and serve data, and you'll prepare and govern the datasets that AI models and AI agents depend on to behave safely and correctly. You'll also use AI-assisted engineering to accelerate your own delivery, while keeping clear human ownership of every outcome.
You'll be part of the Data Platforms team within the Insights and Data Global Practice, which has seen strong, sustained growth across a wide range of sectors. Data Platforms is home to Data Engineers, Platform Engineers, Solution Architects and Business Analysts driving our customers' digital and data transformation on modern cloud platforms. We specialise in the latest frameworks, reference architectures and technologies across AWS, Azure and GCP, and data platforms such as Databricks and Snowflake.
PLEASE NOTE:
Security Clearance: To be successfully appointed to this role, must be eligible to obtain Security Check (SC)clearance or DV (Developed Vetting clearance).
To obtain SC clearance, the successful applicant must have resided continuously within the United Kingdom for the last 5 years, along with other criteria and requirements.
Throughout the recruitment process, you will be asked questions about your security clearance eligibility such as, but not limited to, country of residence and nationality.
Some posts are restricted to sole UK Nationals for security reasons; therefore, you may be asked about your citizenship in the application process.
The Focus of Your Role
You'll build and run the data products that modern AI depends on. That means classic, high-quality data engineering (batch and streaming pipelines, lakehouse and warehouse patterns, well-governed and auditable datasets), and the newer discipline of engineering data for AI and agents: retrieval pipelines, embeddings and vector stores, feature and context preparation, and the observability needed when the systems consuming your data are non-deterministic and can't simply be unit tested.
You'll build and run the data products that modern AI depends on. That means classic, high-quality data engineering (batch and streaming pipelines, lakehouse and warehouse patterns, well-governed and auditable datasets), and the newer discipline of engineering data for AI and agents: retrieval pipelines, embeddings and vector stores, feature and context preparation, and the observability needed when the systems consuming your data are non-deterministic and can't simply be unit tested.
What you’ll be doing:
- You'll work within agreed security and compliance boundaries throughout, which matters in the regulated and public sector environments we operate in.
- Build and maintain data pipelines and data products. Use appropriate ETL/ELT and distributed processing to ingest, transform and curate trusted data in cloud storage and analytics platforms.
- Engineer data for AI and agentic use cases. Prepare curated, governed datasets and features for analytics and ML; build retrieval and embedding pipelines (vector stores, chunking, metadata) that serve RAG and agent workflows; and help expose data to AI systems through patterns such as APIs and the Model Context Protocol (MCP).
- Apply data modelling, quality and governance. Develop scalable models, apply validation and quality checks, and maintain lineage and documentation so data products are reliable and auditable. This matters even more when an AI agent, not just a human, is acting on them.
- Build in observability for non-deterministic systems. Implement logging, alerting and SLAs for production pipelines, and contribute to evaluation and monitoring of AI-facing data and outputs (e.g. drift, retrieval quality, anomaly detection) with clear human review.
- Use AI-assisted engineering responsibly. Accelerate development, testing and documentation with coding assistants, applying good prompt hygiene, protecting confidential data, and owning the results.
- Collaborate across teams. Work with business analysts, platform engineers, data scientists and DevOps to deliver secure, well-tested solutions in an agile environment, communicating your design decisions clearly.
- Keep raising the bar. Use version control, code review, automated testing and CI/CD; share knowledge, contribute to accelerators and standards, and keep current with modern data and AI engineering practice.
What You Will Bring and Experience Needed
You'll bring solid, hands-on experience delivering data engineering in production, and a genuine interest in how AI and agentic systems change what good data engineering looks like. You don't need to have done all of the AI-specific work below already, but you should be eager to, and able to show the engineering fundamentals that make it possible.
Essential:
- Hands-on experience delivering data pipelines and data platforms in production environments.
- Strong Python and SQL, with sound software engineering practice (Git, code review, unit testing, CI/CD) and the ability to troubleshoot production issues.
- Experience with distributed data processing and modern lakehouse/warehouse patterns, with good data modelling and performance instincts.
- Experience with cloud data services (storage, compute and orchestration) on at least one major cloud.
- Practical use of AI coding assistants in a real engineering workflow, with an understanding of data confidentiality and secure, responsible use.
- A continuous-learning mindset and clear communication with technical and non-technical colleagues.
Nice to have
- Exposure to GenAI/agentic building blocks: RAG, embeddings and vector search, LLM orchestration (e.g. LangGraph, LlamaIndex, Semantic Kernel), or LLM evaluation/observability (e.g. LangSmith, Ragas).
- Streaming and event-driven experience (e.g. Kafka, Spark Structured Streaming).
- Relevant cloud and/or data engineering certifications.
Specialist tracks (we'd love depth in one of these; you don't need all of them):
- Azure / Databricks Azure Data Lake Storage, Databricks, Apache Spark, Delta Lake, Azure Data Factory, MLflow; and, for the AI edge, Databricks Mosaic AI, Vector Search and Unity Catalog, or Azure OpenAI, Azure AI Foundry and Azure AI Search.
- AWS: Glue, Lambda, Step Functions, Kinesis, EMR, Athena, Redshift, S3 data lakes; and, for the AI edge, Amazon Bedrock, Knowledge Bases for Bedrock
Additional Info
Hybrid working: The places that you work from day to day will vary according to your role, your needs, and those of the business; it will be a blend of Company offices, client sites, and your home; noting that you will be unable to work at home 100% of the time.
If you are successfully offered this position, you will go through a series of pre-employment checks, including identity, nationality (single or dual) or immigration status, employment history going back 3 continuous years, and unspent criminal record check (known as Disclosure and Barring Service)
What we’ll offer you
You will be encouraged to have a positive work-life balance. Our hybrid-first way of working means we embed hybrid working in all that we do and make flexible working arrangements the day-to-day reality for our people. All UK employees are eligible to request flexible working arrangements.
You will be empowered to explore, innovate, and progress. You will benefit from Capgemini’s ‘learning for life’ mindset, meaning you will have countless training and development opportunities from thinktanks to hackathons, and access to 250,000 courses with numerous external certifications from AWS, Microsoft, Harvard Manage Mentor, Cybersecurity qualifications and much more.
Why we’re different
At Capgemini, we help organisations across the world become more agile, more competitive, and more successful. Smart, tailored, often ground-breaking technical solutions to complex problems are the norm. But so, too, is a culture that’s as collaborative as it is forward thinking. Working closely with each other, and with our clients, we get under the skin of businesses and to the heart of their goals. You will too.
Capgemini is proud to represent nearly 130 nationalities and its cultural diversity. Our holistic definition of diversity extends beyond gender, gender identity, sexual orientation, disability, ethnicity, race, age, and religion. Capgemini views diversity as everything that makes us who we are as an organization, including our social background, our experiences in life and work, our communication styles and even our personality. These dimensions contribute to the type of diversity we value the most: diversity of thought.
London, GB Bristol, GB Birmingham, GB Newcastle upon Tyne, GB Manchester, GB