Mid-Level Data Engineer

  • Full-time

Company Description

Kaelo provides essential healthcare solutions ensuring the physical and psychosocial wellbeing of all South Africans towards lasting social change. Kaelo meets the Healthcare needs of Corporate and Retail clients across South Africa – products offerings include Medical Insurance, Medical Aid, Gap Insurance, Kaelo Money and employee assistance programmes.

Job Description

Seeking a talented and results-oriented Mid-Level Data Engineer to join our growing team. In this pivotal role, you will play a key role in designing, developing, and maintaining our data infrastructure, ensuring the efficient and reliable flow of data to support critical business needs. You will collaborate with data analysts, data scientists, and other stakeholders to translate business requirements into robust data solutions, ultimately driving data-driven decision making across the organization.

Job requirements:

Role Purpose

The Data Engineer designs, builds, and maintains the data pipelines, platforms, and models that enable reliable, secure, and scalable data flow across the organisation. Reporting to the Head of Data, the incumbent translates the data strategy into hands-on engineering delivery — building modern cloud data platforms, implementing layered (medallion) data architectures, and developing semantic layers that make trusted data easily consumable for reporting, analytics, and AI initiatives.

The role combines strong engineering discipline with a practical, business-outcome focus — ensuring that data is accurate, well-modelled, well-governed at the engineering level, and readily accessible to analysts, data scientists, and business stakeholders.

Key Responsibilities

1. Data Architecture & Cloud Platform Engineering

  • Design, build, and maintain scalable cloud data platforms using Google BigQuery and/or Microsoft Fabric.
  • Implement and maintain a medallion (bronze/silver/gold) data architecture, ensuring clear separation between raw, cleansed/conformed, and business-ready data layers.
  • Develop and optimise data warehouse and lakehouse structures, including schemas, partitioning, clustering, and storage strategies for cost and performance.
  • Contribute to enterprise data architecture standards and support migration of legacy data sources into modern cloud platforms.

2. Data Pipeline Development & Integration

  • Design, build, and manage robust ELT/ETL pipelines for batch and real-time data processing using industry-standard tools and frameworks.
  • Integrate data from multiple internal and external systems into trusted enterprise data repositories.
  • Automate data ingestion, transformation, and orchestration workflows to reduce manual effort and improve reliability.
  • Apply strong programming skills (Python, SQL) and, where applicable, distributed processing frameworks (e.g. Spark) to build efficient pipelines.

3. Semantic Layer & Business Intelligence Enablement

  • Design, build, and maintain a semantic layer that translates raw and modelled data into consistent, business-friendly metrics, dimensions, and definitions.
  • Ensure a single source of truth for key business metrics, reducing conflicting definitions across reports and tools.
  • Partner with BI, analytics, and data science teams to expose well-modelled datasets through the semantic layer for reporting, dashboarding, and self-service analytics.

4. Data Quality, Testing & Monitoring

  • Develop and implement automated data quality checks, validation rules, and testing across pipelines and data layers.
  • Build proactive monitoring, alerting, and logging to identify and resolve data issues before they impact the business.
  • Troubleshoot data-related incidents, perform root cause analysis, and implement preventative fixes.
  • Maintain version control, CI/CD pipelines, and testing practices for data engineering code.

5. Data Security, Access & Documentation

  • Implement secure data storage and handling practices, applying appropriate access controls at the platform and dataset level.
  • Support compliance with relevant data protection and governance requirements (e.g. POPIA) within engineering processes.
  • Create and maintain clear technical documentation for pipelines, data models, and the semantic layer to support knowledge transfer.
  • Work closely with data analysts, data scientists, and the Head of Data to translate business and analytical requirements into technical specifications.
  • Stay current on emerging data engineering tools and practices, recommending improvements to the organisation's data infrastructure.

6. AI Readiness & Advanced Analytics Enablement

  • Ensure data pipelines, data models, and the semantic layer are AI-ready — well-structured, high-quality, and easily consumable by machine learning and AI systems.
  • Prepare and maintain trusted, feature-ready datasets that support machine learning model development, training, and deployment.
  • Enable AI-based data analytics capabilities — such as predictive analytics, natural language querying, and generative AI/BI copilots — by ensuring underlying data is accurate, well-governed, and readily accessible.
  • Collaborate with data scientists and analytics teams to support AI/ML use cases, from data preparation through to production deployment.
  • Stay abreast of emerging AI-driven data engineering practices (e.g. vector databases, retrieval-augmented generation pipelines) and assess their application to the organisation's data platform.

 

Qualifications

Qualification:

 

  • Bachelor's degree in Computer Science, Information Systems, Data Science, Engineering, Mathematics, or a related field.
  • Relevant cloud platform certification (Google Cloud, Microsoft Azure/Fabric, or AWS) advantageous.

Experience

Essential

  • Minimum 4–6 years' experience as a Data Engineer or in a similar data platform/engineering role.
  • Hands-on experience with Google BigQuery and/or Microsoft Fabric (or an equivalent modern cloud data platform).
  • Demonstrated experience implementing a medallion (layered) data architecture in a production environment.
  • Demonstrated experience designing and implementing a semantic layer for reporting and analytics consumption.
  • Proven experience designing, developing, and deploying ELT/ETL data pipelines.
  • Experience with data warehousing, dimensional modelling, and data quality practices.

Preferred

  • Experience with orchestration tools (e.g. Airflow, Fabric Data Factory/pipelines, dbt) and version control (Git) with CI/CD for data workflows.
  • Exposure to enterprise BI tools (e.g. Power BI, Looker) and enabling data for AI/ML use cases.
  • Exposure to AI-based data analytics capabilities, such as predictive analytics, generative AI/BI copilots, or ML feature engineering.
  • Experience operating within a regulated industry such as financial services, healthcare, or health insurance.

Technical Competencies

  • Google BigQuery and/or Microsoft Fabric
  • Medallion (Bronze/Silver/Gold) Data Architecture
  • Semantic Layer Design & Business Metrics Modelling
  • ETL/ELT Development & Pipeline Orchestration
  • Advanced SQL and Python
  • Distributed Data Processing (e.g. Spark)
  • Dimensional Data Modelling & Data Warehousing
  • Data Quality, Testing & Observability
  • Version Control & CI/CD for Data Pipelines
  • Cloud Data Security & Access Management
  • Business Intelligence & Self-Service Analytics Enablement
  • AI-Ready Data Preparation & Enablement
  • AI-Based Data Analytics (predictive analytics, GenAI/BI copilots)

Core Competencies

  • Problem Solving & Technical Judgement — resolves complex data engineering challenges with sound, practical solutions.
  • Ownership & Accountability — takes responsibility for the reliability and quality of data delivered.
  • Collaboration — works effectively with analysts, data scientists, and business stakeholders.
  • Attention to Detail — maintains high standards of data accuracy, consistency, and documentation.
  • Continuous Improvement — proactively seeks opportunities to improve platforms, pipelines, and processes.

Key Performance Indicators (KPIs)

  • Pipeline reliability, uptime, and processing success rates.
  • Data quality and integrity scores across bronze, silver, and gold layers.
  • Timely delivery of engineering tasks aligned to the data roadmap.
  • Adoption and accuracy of the semantic layer for reporting and analytics.
  • Reduction in manual data processes through automation.
  • Readiness and usability of data for AI, machine learning, and advanced analytics initiatives.

Knowledge/Experience:

  • Minimum of  3 years of experience as a Data Engineer or similar role.
  • Proven experience in designing, developing, and deploying data pipelines.
  • Three or more years of experience with Python, SQL, PostgreSQL and data visualization/exploration tools such as Power BI or Looker
  • Experience with data warehousing and data modelling concepts.
  • Excellent problem-solving and analytical skills with a data-driven approach.
  • Ability to work on a dynamic, analytical-oriented team that can work on concurrent projects.

 

Additional Information

Personal Attributes

  • Excellent problem-solving skills backed by solid technical knowledge.
  • Excellent customer service and communication skills, with the ability to explain technical concepts to non-technical users.
  • Excellent organisational and time-management skills, with the ability to manage multiple priorities in a fast-paced environment.
  • A versatile and service-oriented mind-set.
  • Good communication skills.
  • Understands the importance of documentation.
  • Open to learn, but also willing to teach.

 

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply