Faster chat, better deals — Get the App

Data Engineer - Streaming & Batch Pipelines

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience level1 to 2 years
Education levelBachelor's Degree

Description

**REQUIREMENTS** Experience * 3+ years of experience as a Data Engineer in high-volume production environments, with solid expertise in Apache Spark for both batch processing and Spark Structured Streaming, using PySpark or Scala. Education * Technical degree-level education: Bachelor's degree in Engineering, Technology Sciences, Mathematics, Economics, Physics, Chemistry, Statistics, or similar. Technical skills * Apache Spark with PySpark or Scala: batch and Spark Structured Streaming. * Delta Lake: MERGE, upserts, time travel, file optimization, and schema management. * Apache Kafka and Confluent: producers and consumers, partitioning, Schema Registry, and CDC via Debezium. * AWS data ecosystem: Glue, Data Catalog, ETL, Lambda, S3, EMR, and EMR Serverless. * AWS Lake Formation: table- and column-level permissions and access governance. * Amazon SageMaker: integration of data pipelines with machine learning workflows. * Advanced SQL and dimensional data modeling. * Change Data Capture and Slowly Changing Dimensions. * Git, CI/CD, and best practices for testing in data projects. Additional information Preferred qualifications: * Experience with Kafka Streams or ksqlDB. * Experience with large-scale lookup table architectures, including point lookup, bucketing, and Bloom Filters. * Knowledge of Terraform or CloudFormation for infrastructure-as-code. * Experience in regulated industries, especially financial services or payment processing. * AWS certifications such as Data Analytics Specialty or Solutions Architect. * We seek a candidate capable of performing deep debugging in distributed systems, demonstrating autonomy and a "you build it, you run it" mindset. **KEY RESPONSIBILITIES** * Design, build, and operate real-time and batch data ingestion and transformation pipelines on a native AWS streaming platform based on Kafka/Confluent, Spark Structured Streaming, and Delta Lake, with focus on reliability, performance, and scalability. Main activities * Design, develop, and maintain streaming pipelines with Spark Structured Streaming and batch processes on AWS. * Implement CDC logic, enrichment joins, checkpoint and offset management, and DLQ handling. * Integrate and maintain Confluent Kafka connectors and components, including topics, Schema Registry, partitioning, and throughput optimization. * Build and optimize Delta Lake tables using partitioning, Z-Ordering or Liquid Clustering, compaction, MERGE/upsert operations, and schema management. * Develop and maintain AWS Glue jobs, including crawlers, catalog, and ETL processes, as well as AWS Lambda functions for orchestration and event-driven processing. * Operate and optimize Spark workloads on EMR and EMR Serverless through memory tuning, partitioning, shuffle optimization, Adaptive Query Execution, and bottleneck resolution. * Collaborate with Data Science and Machine Learning teams to prepare datasets and feature pipelines for SageMaker. * Design strategies for managing large-scale reference or lookup tables, using point lookup, caching, and partitioning to perform efficient enrichment joins. * Implement monitoring, alerting, and resilience mechanisms for 24/7 pipelines, including heartbeats, retries, and checkpointing. * Evolve the internal YAML- and Spark-based ETL framework by applying CI/CD, testing, and infrastructure-as-code.

Source: indeed
Some content was automatically translated

Posted by

David Muñoz

Indeed · HR

Location

David Muñoz

Indeed · HR

Similar jobs

Data Engineer - Streaming & Batch Pipelines job by Indeed in 2026 | ok.com