Apply now »

AI Harness & Evaluations Engineers

 

About the job you’re considering

Hybrid working: The places that you work from day to day will vary according to your role, your needs, and those of the business; it will be a blend of Company offices, client sites, and your home; noting that you will be unable to work at home 100% of the time.

If you are successfully offered this position, you will go through a series of pre-employment checks, including, identity, nationality (single or dual) or immigration status, employment history going back 3 continuous years, and unspent criminal record check (known as Disclosure and Barring Service)

 

About us

We are developing a new AI-native product organisation within Capgemini Financial Services. We build products, not projects: software for insurance claims, payment operations, and health operations, sold to banks, insurers, and health plans. Three product lines run on one shared platform, built by a deliberately small, senior team. Our engineering model is agentic: engineers author the specifications, tooling, evaluation suites, and guardrails, and AI agents do most of the implementation. Humans own every consequential decision, and in our regulated domains some decisions are human-only by design.

The role

In a regulated domain, an AI product is only as trustworthy as the evidence that it works. You will own that evidence for one product line: the golden datasets that define correct behaviour, the gates that stop a regression reaching a client, the drift monitors that catch quality decaying in production, and the evidence pack a bank or insurer's model-risk function reads before they let us anywhere near their operations. You will build those datasets alongside a full-time domain expert, because in claims and payments the definition of a correct answer is a professional judgment, not a label.

You also own how your line runs its harness: the steering files, specification conventions, and agent tooling your engineers work inside, adapted from the platform group's patterns to the shape of your domain.

What you will own

  • The line's evaluation suites: domain golden datasets that are human-anchored and domain-anchored, never derived from the system under test
  • Gate design: per-cohort thresholds, statistical gating with honest sample sizes, and judge calibration with pinned versions so a gate means the same thing this month as last
  • Evaluations wired into CI as merge gates, and into production as drift monitors that auto-score a sampled share of live traffic
  • The failure flywheel: production failures mined into permanent regression cases, so the same class of error cannot recur silently
  • The evaluation evidence pack that services customer model-risk validation (SR 11-7 vendor obligations), written to be read by people outside engineering
  • The line's local harness: steering files, specification constitutions, and agent tooling, kept current with platform patterns and pruned when model progress makes a guard unnecessary

What you will need

  • Machine learning or data engineering experience with genuine measurement rigour, and production systems you have measured rather than only built
  • Statistics you can defend: cohort design, sample sizing, confidence intervals, and the discipline to say a result is not yet real
  • Hands-on evaluation engineering for language models: golden datasets, LLM-as-judge calibration, regression gates in CI, and the failure modes of each
  • Strong Python, and comfort operating CI/CD systems
  • The willingness to build domain literacy in the seat, working daily with a claims or payments expert
  • Daily, hands-on use of AI coding assistants as part of your own workflow

What sets you apart

  • You have produced evaluation evidence that an external risk, audit, or regulatory function accepted
  • Adversarial or red-team dataset construction, and an understanding of where that work stops and safety testing starts
  • You have run a human annotation programme: recruiting, calibrating, and measuring agreement between expert annotators
  • Reinforcement learning or fine-tuning pipelines fed by evaluation verdicts and production traces
  • Published or open-source work on evaluation methodology or tooling

The reference stack

The reference technology stack for this role is our supported paved road: self-hosted LangSmith and LangGraph Platform as the agent runtime and evaluation plane, model providers behind a swappable gateway seam, PostgreSQL with pgvector plus ClickHouse and S3-compatible object storage as the data platform, Neo4j Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, OpenTelemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyperscaler or on-prem. A tool-for-tool match is not expected: analogous experience counts fully. If you have built and operated systems of this shape on comparable components (a different orchestration framework, graph engine, evaluation platform, or serving stack), you have what we are looking for.

How we work

  • Engineers write specs, harnesses, evals, and guardrails; AI agents execute the implementation loops. Review, not typing, is where engineering judgment goes.
  • Three human gates govern everything we ship: spec approval, merge, and release. Regulated code paths (money movement, authentication, cryptography, secrets) are always human-owned.
  • Small and senior by design. No scrum masters and no separate delivery-management layer; quality is owned inside the product team, by the engineers who build and the quality and evaluation engineers who work alongside them.
  • Domain experts (claims practitioners and payment scheme experts) are full-time members of your team, not advisors you consult.

Success in year one

  • No regression reaches a client that your suite could have caught, and you can name the ones it stopped
  • The evaluation evidence pack passes a bank or insurer model-risk review on first submission
  • Every recurring production failure from the last two quarters exists as a permanent regression case
  • Gate thresholds are set from evidence rather than intuition, and the team trusts a green gate enough to release on it 

We are a Disability Confident Employer

Capgemini is proud to be a Disability Confident Employer (Level 2) under the UK Government’s Disability Confident scheme. As part of our commitment to inclusive recruitment, we will offer an interview to all candidates who:

  • Declare they have a disability, and 
  • Meet the minimum essential criteria for the role.

Please opt in during the application process.

 

Make it real – what does it mean for you?

  • We realise a Total Reward package should be more than just compensation.  At Capgemini we offer range of core and flexible benefits and have a Peer Recognition Portal called Applaud.
  • You’d be joining an accredited Great Place to work for Wellbeing in 2024. Employee wellbeing is vitally important to us as an organisation.  We see a healthy and happy workforce a critical component for us to achieve our organisational ambitions. 
    To help support wellbeing we have trained ‘Mental Health Champions’ across each of our business areas, and we have invested in wellbeing apps such as Thrive and Peppy.
  • You will be empowered to explore, innovate, and progress. You will benefit from Capgemini’s ‘learning for life’ mindset, meaning you will have countless training and development opportunities from thinktanks to hackathons, and access to 250,000 courses with numerous external certifications from AWS, Microsoft, Harvard ManageMentor, Cybersecurity qualifications and much more.

Capgemini. Make it real.

 

Why you should consider Capgemini

Growing clients’ businesses while building a more sustainable, more inclusive future is a tough ask.  When you join Capgemini, you’ll join a thriving company and become part of a collective of free-thinkers, entrepreneurs and industry experts.  We find new ways technology can help us reimagine what’s possible.  It’s why, together, we seek out opportunities that will transform the world’s leading businesses, and it’s how you’ll gain the experiences and connections you need to shape your future.  By learning from each other every day, sharing knowledge, and always pushing yourself to do better, you’ll build the skills you want. You’ll use your skills to help our clients leverage technology to innovate and grow their business. So, it might not always be easy, but making the world a better place rarely is.

 

About Capgemini

Capgemini is an AI-powered global business and technology transformation partner, delivering tangible business value. We imagine the future of organisations and make it real with AI, technology and people. With our strong heritage of nearly 60 years, we are a responsible and diverse group of 420,000 team members in more than 50 countries. We deliver end-to-end services and solutions with our deep industry expertise and strong partner ecosystem, leveraging our capabilities across strategy, technology, design, engineering and business operations. The Group reported 2024 global revenues of €22.1 billion.


Make it real 

Ref. code:  529201
Posted on:  21 Aug 2026
Experience Level:  Experienced Professionals
Contract Type:  Permanent
Location: 

London, GB

Brand:  Capgemini
Professional Community:  Software Engineering

Apply now »