Description
**REQUIREMENTS**
Experience
* 3\+ years of experience in data governance, data quality, or data architecture, with a solid technical engineering foundation. This position requires hands-on experience implementing technical controls and automation—not limited to process or documentation tasks.
Education
* Technical degree-level education: Bachelor's degree in Engineering, Technology Sciences, Mathematics, Economics, Physics, Chemistry, Statistics, or similar.
Technical skills
* AWS Lake Formation and AWS Glue Data Catalog.
* Data cataloging, crawlers, and classifiers.
* Apache Spark for batch and streaming processing.
* Delta Lake to implement data quality and governance controls.
* Confluent Kafka, Schema Registry, data contracts, and topic governance.
* AWS Lambda to automate event-driven governance controls and workflows.
* Amazon SageMaker and its relationship with data governance for Machine Learning.
* Feature stores and dataset traceability.
* Knowledge of regulatory frameworks such as GDPR, PCI-DSS, or other applicable standards.
* Ability to translate regulatory requirements into technical controls.
* Advanced SQL and ability to audit and validate transformations in data pipelines.
Additional Information
Preferred qualifications:
* Experience with cataloging and lineage tools such as DataHub, Collibra, Amundsen, or OpenLineage.
* Familiarity with data mesh architectures.
* Experience with data domains, data contracts, and data products.
* Experience in regulated industries, especially financial services or payment processing.
* AWS certifications related to security or data analytics.
* We seek a hybrid profile bridging policy and code—rigorous, with strong communication skills to collaborate effectively with Compliance, Legal, Engineering, and Business teams.
**KEY RESPONSIBILITIES**
* Implement and operate data governance—cataloging, quality, security, compliance, and lineage—directly on the AWS-based engineering platform (Confluent, Glue, Spark, Delta Lake), going beyond purely document-based definitions.
Key activities:
* Design and implement the access governance model using AWS Lake Formation, including table-, column-, and row-level permissions, Tag-Based Access Control, and IAM integration.
* Administer and evolve the AWS Glue Data Catalog, managing schemas, classifications, ownership, and data sensitivity—including PII identification.
* Define and implement automated data quality controls over Spark and Delta Lake pipelines, including schema validation, completeness checks, duplicate detection, and anomaly identification.
* Integrate quality and governance controls into AWS Glue, Lambda, and EMR.
* Design retention policies, versioning, and lineage for Delta Lake tables using time travel, Change Data Feed, and change auditing.
* Define, together with Data Engineering, naming conventions, partitioning standards, and data contracts for Kafka topics and Delta Lake tables.
* Establish mechanisms for classification and protection of sensitive data—including encryption, masking, and role- or domain-specific access controls.
* Provide technical and regulatory support to Data Science and Machine Learning teams for governed access to feature stores and certified datasets.
* Audit compliance with controls in pipelines built using Lambda, Glue, and EMR.
* Produce data quality and governance metrics and reports—including data quality scorecards, domain catalogs, and lineage health indicators.
* Serve as the primary liaison between Engineering, Security, Legal, Compliance, and Business units.