Description
**REQUIREMENTS**
Experience
* 3+ years of experience as a Data Engineer in high-volume production environments, with solid expertise in Apache Spark for both batch processing and Spark Structured Streaming, using PySpark or Scala.
Education
* Technical degree-level education: Bachelor's degree in Engineering, Technology Sciences, Mathematics, Economics, Physics, Chemistry, Statistics, or similar.
Technical skills
* Apache Spark with PySpark or Scala: batch and Spark Structured Streaming.
* Delta Lake: MERGE, upserts, time travel, file optimization, and schema management.
* Apache Kafka and Confluent: producers and consumers, partitioning, Schema Registry, and CDC via Debezium.
* AWS data ecosystem: Glue, Data Catalog, ETL, Lambda, S3, EMR, and EMR Serverless.
* AWS Lake Formation: table- and column-level permissions and access governance.
* Amazon SageMaker: integration of data pipelines with machine learning workflows.
* Advanced SQL and dimensional data modeling.
* Change Data Capture and Slowly Changing Dimensions.
* Git, CI/CD, and best practices for testing in data projects.
Additional information
Preferred qualifications:
* Experience with Kafka Streams or ksqlDB.
* Experience with large-scale lookup table architectures, including point lookup, bucketing, and Bloom Filters.
* Knowledge of Terraform or CloudFormation for infrastructure-as-code.
* Experience in regulated industries, especially financial services or payment processing.
* AWS certifications such as Data Analytics Specialty or Solutions Architect.
* We seek a candidate capable of performing deep debugging in distributed systems, demonstrating autonomy and a "you build it, you run it" mindset.
**KEY RESPONSIBILITIES**
* Design, build, and operate real-time and batch data ingestion and transformation pipelines on a native AWS streaming platform based on Kafka/Confluent, Spark Structured Streaming, and Delta Lake, with focus on reliability, performance, and scalability.
Main activities
* Design, develop, and maintain streaming pipelines with Spark Structured Streaming and batch processes on AWS.
* Implement CDC logic, enrichment joins, checkpoint and offset management, and DLQ handling.
* Integrate and maintain Confluent Kafka connectors and components, including topics, Schema Registry, partitioning, and throughput optimization.
* Build and optimize Delta Lake tables using partitioning, Z-Ordering or Liquid Clustering, compaction, MERGE/upsert operations, and schema management.
* Develop and maintain AWS Glue jobs, including crawlers, catalog, and ETL processes, as well as AWS Lambda functions for orchestration and event-driven processing.
* Operate and optimize Spark workloads on EMR and EMR Serverless through memory tuning, partitioning, shuffle optimization, Adaptive Query Execution, and bottleneck resolution.
* Collaborate with Data Science and Machine Learning teams to prepare datasets and feature pipelines for SageMaker.
* Design strategies for managing large-scale reference or lookup tables, using point lookup, caching, and partitioning to perform efficient enrichment joins.
* Implement monitoring, alerting, and resilience mechanisms for 24/7 pipelines, including heartbeats, retries, and checkpointing.
* Evolve the internal YAML- and Spark-based ETL framework by applying CI/CD, testing, and infrastructure-as-code.