Topic
Apache Spark
Spark sits under the largest systems described here: streaming and batch processing of network telemetry at Vodacom, telematics and risk workloads on Databricks in insurance, and Ubunye Engine, an open source framework that runs the same Spark pipeline folder on a laptop, Docker, Kubernetes, object storage, spark-submit and Databricks.
3 work · Wikidata Q7573619
Work
- Generator optimisation and streaming at VodacomReal time analytics and optimisation for a national telecoms network: generator dispatch across 15,000+ sites and tens of millions of events a day.
- Insurance data science: telematics, flood risk and MLOpsLeading insurance data science at ABSA Insurance: telematics processing cut from months to under a day, flood risk across 230,000+ properties, MLOps.
- Ubunye Engine: portable Spark pipelines for data and MLUbunye Engine is an open source Python framework for config driven Spark pipelines. The same task folder runs on a laptop, Docker, Kubernetes or Databricks.
Related topics
Subjects that share work with this one. Those with their own page are linked.
- Data engineering
- MLOps
- Python
- Databricks
- Docker
- Kubernetes
- Technical leadership
- Apache Flink
- Apache Kafka
- Climate risk
- Decision support systems
- Extract, transform, load