
Data Engineer
Компания скрыта (Fintech)
5 000 — 7 000 $/month gross
📍 Worldwide
Remote
Position
Data Engineering
Seniority level
Senior
English
B2 — Upper-Intermediate
Experience
5+ years
Technologies / Tools
Apache Airflow
Apache Spark
Python
Polars
PySpark
SQL
Apache Arrow
Docker
Kubernetes
Apache Iceberg
The Data Engineering team is responsible for designing, building, and maintaining the Market Data Platform — a lakehouse infrastructure spanning the full path from raw exchange feeds to reliable, petabyte-scale data for research, backtesting, and real-time trading.
Key Responsibilities
- Capture & Ingestion. Own the full capture path from wire to lake: decode and normalize raw exchange feeds (PCAP, multicast UDP, ITCH, FIX) and vendor sources (OneTick, Refinitiv, Bloomberg, ICE) into a unified canonical model with nanosecond timestamps. Build batch and streaming pipelines (Airflow, Spark, dbt) for tick and reference data. Own L2/L3 order-book reconstruction with gap handling. Provide Python and Rust producer SDKs for internal feed handlers.
- Storage & Modeling — Apache Iceberg. Own the Iceberg-over-S3 lakehouse: design partitioning, sort orders, and row-group layouts for fast scans; manage schema evolution, snapshots, time travel, compaction, and TTL. Maintain reference data as slowly changing tables with point-in-time correctness for backtests. Drive storage cost optimization through compaction, tiering, and snapshot expiry.
- Tooling & Libraries. Build libraries for schema management, data contracts, validation, and lineage on top of the Iceberg catalog. Develop shared access services (Spark + Polars) so research, backtesting, and trading share a single normalized data layer, including gap detection and PCAP-vs-lake reconciliation.
- Reliability & Observability. Embed monitoring, alerting, SLAs/SLOs, and CI/CD across capture and pipeline layers on Kubernetes (EKS). Own data-quality dashboards and incident runbooks for the capture fleet.
- Collaboration. Partner with Quant Research, Data Science, Backend, and DevOps to translate requirements into platform capabilities and champion market-data engineering best practices.
Qualifications
- 5+ years of experience building production-grade data systems, with deep knowledge and understanding of the problems that come with running data lakes/lakehouses at scale: data layout and partitioning strategy, ingestion and backfills, small-file and metadata growth, consistency and late-arriving data, query performance, storage cost, and operational failure modes.
- Hands-on experience with Apache Iceberg (or comparable table formats such as Delta/Hudi): partitioning, schema evolution, snapshots, compaction, and catalog operations; familiarity with Apache Arrow for zero-copy, columnar in-memory interchange.
- Expert-level Python (incl. Polars and/or PySpark).
- Modern orchestration (Airflow) and distributed processing/query engines (Apache Spark, StarRocks).
- Advanced SQL: complex aggregations, window functions, query optimization, partition pruning.
- Solid fundamentals in Linux, containerization (Docker, Kubernetes/EKS), and cloud object storage (AWS S3).
- DevOps & observability: CI/CD, infrastructure-as-code (Terraform), GitOps (ArgoCD), and metrics/dashboards/alerting (Grafana, Prometheus).
- Strong grasp of structured + unstructured/binary data and storage optimization: partitioning, compression, cost management.
- English fluency (B2+) for documentation and collaboration in an international team.
Nice to have
- Experience with market data and/or network packet capture: decoding PCAP, exchange feed protocols (ITCH, FIX/FAST, multicast UDP), order-book reconstruction, and time-series at scale. Willingness to learn this domain is expected.
- Experience normalizing market data from multiple vendors (OneTick, Refinitiv/Reuters, Bloomberg, ICE) into a unified schema and symbology.
- Rust, relevant for high-performance capture/decoding.
What we offer
- Fully remote setup with a daily overlap window from 12:00 to 18:00 GMT+3.
- Reimbursement for health insurance, sports, and personal development.
- A genuinely flat structure with real autonomy over how you work.
- Room to grow the role as the team scales.

About company Компания скрыта (Fintech)
Industry
Финтех
Company size
201 - 500
The company name is under an NDA. A proven Fintech global company with an advanced stack and decade of successful trading experience is developing a proprietary platform to trade thousands of instruments across dozens of markets. The recruiter will disclose all details in person immediately upon response.