Remote job
Remote confirmedData Engineer
payabl.
Our assessment
- Our reading of the full posting text confirms it: fully remote.
- The posting states no salary. Comparable roles in our index (77 postings): median 5,625 euros per month, middle range 4,535 to 6,875 euros.
- 11 more open roles from this employer in our index. 9 of them fully remote.
This section only: calculated automatically by nomado24, from our own job index and our own reading of the posting text. Not stated by the employer.
Job description
payabl. empowers businesses to grow through payments innovation and banking services. Our ambition is to expand our strong portfolio of global financial services we provide to businesses and make them all available in one place on our platform we call payabl.one. As a licensed financial company with principal membership with card schemes, we specialize in global payments and providing businesses with multi-currency accounts.
The role is about:
Data Team, where you play a vital role in collecting, analyzing, and interpreting data to support decision-making across the organization. Your tasks include data collection, ensuring quality, and developing databases. You're skilled in SQL, Python for data manipulation and visualization tools. Collaborative and communicative, you work with various teams to understand their needs and provide actionable insights. Passionate about staying updated with new technologies, you drive innovation and contribute to a culture of continuous improvement.
Location: remote Poland, Portugal or on-site Cyprus
Reporting to: Head of Engineering
What you will do:
Lakehouse Architecture and Data Platform Design
- Design, build, and maintain scalable data lakehouse solutions on AWS using S3, Apache Iceberg, and AWS Glue Catalog as core platform components.
- Contribute to the evolution of the medallion architecture, ensuring that bronze, silver, and gold layers are reliable, performant, and aligned with business needs.
CDC and Streaming Data Ingestion
- Build and support real-time and near-real-time ingestion pipelines that stream operational data from on-premise databases into AWS using Debezium, Kafka, Kafka Connect, and Iceberg sinks.
- Monitor and troubleshoot streaming pipelines, including connector issues, schema changes, data consistency, and ingestion reliability.
Data Modelling and Business-Ready Layers
- Design and implement silver and gold layer datasets that transform raw operational data into trusted, business-ready data products.
- Work closely with analysts, analytics engineers, and business stakeholders to understand requirements and create reusable, well-structured data models.
Batch and Distributed Processing
- Develop and maintain batch and distributed processing jobs using PySpark on AWS EMR and AWS Glue.
- Optimize data transformation jobs for performance, scalability, reliability, and cost efficiency.
Data Integration and Orchestration
- Build and maintain data workflows using Apache Airflow for API ingestion, batch processing, and orchestration across the medallion architecture.
- Support integrations from external and third-party systems using tools such as Airbyte.
Data Quality, Governance, and Reliability
- Implement data quality checks, validation processes, and reconciliation logic to ensure trusted and consistent datasets across the platform.
- Contribute to data governance practices around documentation, ownership, lineage, access, and compliance.
Infrastructure and DevOps Collaboration
- Collaborate on AWS infrastructure and deployment practices, working with services such as S3, Glue, EMR, IAM, EKS, and related platform components.
- Support infrastructure-as-code workflows using Terraform and Terragrunt in collaboration with engineering and infrastructure teams.
What we need:
- Minimum 3+ years of experience in data engineering or related roles.
- Strong experience with SQL and data modeling for analytics and reporting use cases.
- Strong programming experience with Python, ideally including PySpark.
- Experience designing and maintaining ETL/ELT pipelines in production environments.
- Experience with real-time or near-real-time data ingestion.
- Experience working with Kafka or similar streaming technologies.
- Experience with CDC concepts and tools such as Debezium.
- Experience with data lake or lakehouse architectures on cloud platforms.
- Hands-on experience with AWS data services such as: S3, Glue Catalog, Glue Jobs, EMR, IAM, Related AWS services
- Experience with Apache Iceberg, Delta Lake, or similar open table formats.
- Experience designing curated analytics layers, such as silver and gold layers in a medallion architecture.
- Experience with orchestration tools such as Apache Airflow or Dagster.
- Experience working with relational databases such as MySQL, MariaDB, or PostgreSQL.
- Working knowledge of Unix/Linux environments and shell scripting.
- Understanding of data quality, governance, lineage, and production monitoring concepts.
Good to Have
- Experience with Apache Druid, ClickHouse, Snowflake, Databricks, or other analytical databases and warehouse platforms.
- Experience with Airbyte or similar data integration tools.
- Experience with dbt or collaboration with analytics engineering teams.
- Experience with Terraform and Terragrunt.
- Experience with Docker and Kubernetes.
- Experience optimizing Spark jobs on EMR or AWS Glue.
- Experience with data visualization tools such as: …
This role is provided by an external source. Applications are handled on the source website.
