Building the capability around insurance data science
Good models are only one part of a working data-science function. The data, engineering, governance, tooling and team around them decide whether anything survives in production.

The problem
The function had capable people building good models, but the path from an idea to a governed production system depended on individual knowledge and one-off effort. Historically it leaned more toward BI and analysis than a production data-science engineering operating model, which is a common and reasonable place for an insurance analytics team to start. The question was how to move it toward something that could reliably build, deploy, govern and improve analytical products.
Why it mattered
A data-science capability is more than a collection of models. Underwriting, retention, fraud, telematics and climate risk all need work that can be reproduced, operated and improved by someone other than its author. Without the surrounding engineering and governance, a good model is a one-off.
The context
A large, regulated enterprise, mid-migration from on-premise data and workflows toward cloud. Databricks is the enabling platform rather than the point. Confidentiality limits what can be said about specific systems, so what follows is scope and consequence, not internals.
What I did
I lead the Insurance Data Science capability across four connected layers. Strategy and modernisation: I help define the technical and analytical direction, moving toward cloud-first ways of working, scalable data and ML systems, and governed analytical products. Data and ML engineering: I have been helping introduce hands-on practice across the whole lifecycle, from data engineering and experimentation through deployment, monitoring, governance and maintenance, including CI/CD, model lifecycle management, testing and reproducibility. People and ways of working: I lead a multidisciplinary team of senior data scientists and data engineers, and much of the job is structure rather than technology, clearer ways of working, Agile delivery, engineering standards, documentation, Show and Tell sessions, peer learning and reducing knowledge silos. Applied systems: alongside that I still build. Telematics and behavioural risk supporting products such as Activate; high-resolution geospatial flood and natural-catastrophe risk models, built with XGBoost against JBA ground-truth data, so physical exposure can be understood at property and portfolio level before losses occur; and customer intelligence work that gives the business a richer view of behaviour, value, needs and risk rather than a single recommendation model.
What changed
The core telematics platform was modernised using Ubunye Engine, cutting data pipeline processing latency from about two months to under 24 hours at scale. Enterprise AI governance and CI/CD were institutionalised across the production portfolio using Databricks Asset Bundles, MLflow and Unity Catalog. The effect is less manual intervention, faster data availability, more consistent processing and clearer ownership, which is what lets several analytical products run at once.
Who benefited
Insurance operations and underwriting, through earlier visibility of physical risk and better behavioural understanding; the data scientists on the team, who can ship more reliably and depend less on any one person; and ultimately customers, through more relevant decisions and interactions.
What remained
Reusable engineering practice, governance and CI/CD across a production portfolio, a modernised telematics platform, geospatial risk models across 230,000+ insured properties, hyperpersonalisation processing 2M+ daily telematics signals for retention, next-best-action and customer lifetime value, and a more self-sufficient team.
Technical context
Databricks, Spark and PySpark, Databricks Asset Bundles, MLflow, Unity Catalog, CI/CD, XGBoost, JBA flood ground-truth data, geospatial modelling, MLOps and model lifecycle management, AI governance, cloud migration from on-premise. Ubunye Engine underpins part of the telematics modernisation.
Related
Making practical AI knowledge easier to reach
Public AI conversation is split between hype and inaccessible research, with little plain explanation of what building with these systems actually involves.
WorkNetwork intelligence and optimisation
Enormous volumes of noisy, weakly labelled operational data from a national network had to become decisions about where finite resources and intervention were needed most.
WorkUbunye Engine
Every team rebuilds the same pipeline plumbing, structured differently each time, and the tooling is split across laptops, on-prem clusters and different clouds.
ResearchLong-range seasonal forecasting of 2m-temperature with machine learning
Developed ML models for long-range seasonal temperature forecasting, outperforming traditional numerical weather prediction at extended lead times. Published during IBM Research Africa tenure, integrated into climate intelligence workflows.