Senior Data Engineer - AWS + Spark Required

Remote
Full Time
Costa Rica Studio
Experienced

This remote role is based in Costa Rica and only open to citizens and permanent residents of Costa Rica who do not need visa sponsorship.  

We are seeking an experienced Senior Data Engineer to design, build, and optimize our next-generation transactional data lake and analytics infrastructure. In this role, you will lead the architecture and implementation of scalable batch and streaming pipelines using Apache Spark, Apache Hudi, AWS Glue, and AWS Athena, ensuring high performance, transactional consistency, and low-latency querying for enterprise data.

Responsibilities

  • Data Lake Architecture: Design, implement, and maintain scalable, ACID-compliant data lakehouse architectures utilizing Apache Hudi on AWS S3.
  • Pipeline Development: Build robust, fault-tolerant ETL/ELT pipelines using Apache Spark (PySpark/Scala) and AWS Glue for high-volume batch and real-time streaming data.
  • Query Optimization: Write advanced, performant SQL queries and optimize interactive analytical workloads using AWS Athena.
  • Data Lakehouse Management: Manage Hudi table operations including upserts, soft/hard deletes, compaction, clustering, time-travel queries, and schema evolution.
  • Performance Tuning: Optimize Glue jobs, Spark cluster performance, partitioning strategies, file formats (Parquet/ORC), and query latency across large-scale datasets.
  • Governance & Quality: Enforce data quality check frameworks, metadata management, cataloging via AWS Glue Data Catalog, and data security/privacy standards.
  • Collaboration & Leadership: Mentor junior/mid-level data engineers, participate in code reviews, and partner with Analytics, Data Science, and Product teams to align data modeling with business goals.

Technical Skills

  • Experience: 5+ years of software/data engineering experience, with a track record of architecting distributed data platforms.
  • SQL: Advanced mastery of complex SQL queries, analytical functions, CTEs, and performance tuning techniques.
  • Apache Spark: Deep technical knowledge of Spark architecture (RDDs, DataFrames, Spark SQL, Structured Streaming), memory management, and job optimization.
  • Apache Hudi: Hands-on experience implementing transactional data lakes with Apache Hudi (Copy-on-Write vs. Merge-on-Read, indexing strategies, CDC ingestion, compaction).
  • AWS Glue: Expertise in building serverless Glue ETL pipelines, custom scripts, dynamic frames, Glue Triggers, and managing the AWS Glue Data Catalog.
  • AWS Athena: Strong proficiency with ad-hoc querying, query execution cost/time optimization, dynamic partitioning, and federated queries.
  • Programming: Proficiency in Python or Scala.
  • Data Warehousing & Modeling: Solid grasp of dimensional modeling, data vault, Star/Snowflake schemas, and database internals.

Nice to Have

  • Experience setting up CDC (Change Data Capture) pipelines (e.g., AWS DMS, Debezium) into Apache Hudi
  • Experience with Infrastructure as Code (Terraform, AWS CloudFormation, or AWS CDK)
  • Knowledge of orchestration tools such as Apache Airflow or AWS Step Functions.
  • AWS Certified Data Engineer – Associate or AWS Certified Big Data / Data Analytics Specialty

Strategic Skills 

  • Excellent verbal and written communication skills. 
  • Team player. 
  • Experience working within agile environments. 

This remote role is based in Costa Rica and only open to citizens and permanent residents of Costa Rica who do not need visa sponsorship.  
Share

Apply for this position

Required*
We've received your resume. Click here to update it.
Attach resume as .pdf, .doc, .docx, .odt, .txt, or .rtf (limit 5MB) or Paste resume

Paste your resume here or Attach resume file

Human Check*