Skip to content
View cloudcruncher's full-sized avatar
:octocat:
Working from home
:octocat:
Working from home

Block or report cloudcruncher

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
cloudcruncher/README.md

Hi, I'm Robin 👋

Lead / Senior Data Engineer · AI data engineering · London, UK

I build data platforms for regulated financial services, and lately the data side of agentic AI: getting AI agents safely to the enterprise data they need. 15+ years in software and data engineering, 10+ of them in data engineering, most recently at NatWest Group.

LinkedIn Email


🔭 What I'm working on

  • Agentic AI at NatWest Innovation. I lead the data engineering area for the team's agentic AI projects, connecting agents to API and non-API enterprise systems through MCP servers and direct tool calls, from proof of concept to delivery.
  • Open Lakehouse, my own build of a governed lakehouse for a bank, built around an AI assistant on live customer calls. Interactive tour →
    • CDC streaming (Debezium → Kafka → Spark → Iceberg): a source commit is visible to a colleague in about 10 s
    • One policy for people and agents: Trino + OPA row filters and column masks, and agents act on behalf of the colleague
    • Every AI answer shows how the platform produced it: the Iceberg snapshot it was pinned to, the OPA decision, and the audit row
    • 53 end-to-end checks and a chaos suite (11 components killed and recovered) run on every push in GitHub Actions

📈 Selected impact

Agentic AI Complaints investigation cut from 45–60 min to 2–3 min per case (>95%); due-diligence checks supported on 300+ onboarding applications a day
ESG platform Led the team that built NatWest's ESG and climate platform on Snowflake: 125+ sources, 1 TB+ a month, feeding climate regulatory reporting
Performance 40% faster batch processing (PySpark, dbt) and 30% fewer data-quality incidents through automated monitoring
Enterprise DWH 50% lower ETL load times and 45% fewer data errors on the Lloyds Banking Group warehouse; led a 5-person Data Services team

🧰 Stack

Every day: Python · SQL · Snowflake · Apache Spark / PySpark · Airflow · dbt · PostgreSQL · AWS

Python SQL Snowflake Apache Spark Airflow dbt PostgreSQL AWS

Lakehouse, streaming and governance: Apache Iceberg · Trino · Kafka · Debezium · OPA · Keycloak · Dagster · OpenLineage

Apache Iceberg Trino Apache Kafka Open Policy Agent Dagster

AI data engineering: MCP servers and tool calling · retrieval and context pipelines · Pydantic data contracts · LLM evaluation · AI observability

Delivery: Docker · GitLab CI/CD · GitHub Actions · Terraform · Splunk

Docker GitLab CI GitHub Actions Terraform

🏦 Where I've worked

When Where What
2025 – now NatWest Group, Innovation Senior Data Engineer, agentic systems and LLM deployment. Lead the data engineering area for agentic AI projects
2023 – 2025 NatWest Group, Climate Analytics Led the team delivering the ESG data platform on Snowflake
2021 – 2023 NatWest Group, Retail Analytics Batch pipelines and Snowflake data models for retail decisioning
2016 – 2021 TCS for Lloyds Banking Group Led the Data Services team for the enterprise warehouse (Teradata) and ODS (DB2); built an ML-ready data lake on GCP
2011 – 2016 UST Global, Sopra Steria Mainframe engineering (COBOL, DB2), including for a UK retailer

Domains: retail and commercial banking · financial crime and customer due diligence · complaints · ESG and climate risk · regulatory reporting

Certifications: Big Data on AWS (AWS) · Data Engineering Nanodegree (Udacity) · Google Cloud Engineering (Coursera) · B.Tech Computer Science


The quickest way to reach me is LinkedIn.

Pinned Loading

  1. cloudcruncher cloudcruncher Public

    profile

  2. awesome-opensource-data-engineering awesome-opensource-data-engineering Public

    Forked from gunnarmorling/awesome-opensource-data-engineering

    An Awesome List of Open-Source Data Engineering Projects

  3. data-engineer-handbook data-engineer-handbook Public

    Forked from DataExpert-io/data-engineer-handbook

    This is a repo with links to everything you'd ever want to learn about data engineering