Big data engineering  ·  Contract & freelance  ·  Remote, EU

Data platforms that survive contact with production.

Quarkray is a one-person engineering practice built around large-scale data. Pipelines, lakehouses and streaming systems — scoped, built and handed over by the same senior engineer. No account managers, no junior hand-off, no six-week discovery phase before anyone writes code.

Straight answer on the call, written scope within a few days.

reference pipeline batch + streaming
OLTP / CDC EVENTS SAAS APIS INGEST · CONTRACTS · SCHEMA REGISTRY STREAM · FLINK BATCH · SPARK LAKEHOUSE · ICEBERG ON OBJECT STORAGE BI ML / RAG APPS
Senior only

The engineer who scopes the work is the one who writes it, and the one who hands it over.

Depth + range

Big data platforms as the specialism; APIs, apps and cloud infrastructure when the platform needs them.

Engagement

Fixed-scope build, architecture review, or monthly retainer. Contract or freelance, remote.

Handover

Runbooks, tests, infrastructure as code and a walkthrough — so your team can own it.

01 The usual symptoms

Most data problems don't start as data problems.

They start as a deadline. Something ships that works on Tuesday's data. Then the volume doubles, an upstream schema changes without warning, and the pipeline becomes the thing nobody wants to touch. These four show up in almost every audit.

Symptom 01

The pipeline that fails quietly

A job dies at 3am and a stakeholder finds out first. The missing piece is rarely monitoring — it's idempotent writes, real retries, and a data contract with the upstream team so a renamed column stops being an outage.

Symptom 02

The bill that outgrows the data

Warehouse spend climbing while query patterns stay flat. It's usually file sizes, partitioning and clustering — plus a handful of dashboards quietly scanning full tables on a five-minute refresh.

Symptom 03

Two numbers for one metric

Finance and product disagree because "active customer" is defined in five dashboards and two notebooks. That's a modelling problem with a known fix: one tested definition, in version control, with lineage you can point at.

Symptom 04

The migration stuck at 80%

A lakehouse or warehouse move that has been nearly done for two quarters. The target architecture was designed; the cut-over — dual writes, reconciliation, the order things get switched off — never was.

02 Services

Deep in data. Broad enough to ship the whole thing.

The specialism is big data engineering — platforms that handle real volume without a team of five to babysit them. The range covers everything that platform has to connect to, so there's no hand-off between vendors at exactly the seams where projects break.

Also built when the platform needs it: FastAPI and Node services, React and Next.js front-ends, Terraform, Kubernetes and CI/CD. Full service detail

03 How it works

Three steps, no discovery theatre.

  1. A 30-minute call

    You describe the problem. You get a straight opinion on whether it's worth building, buying, or leaving alone — free, and without a slide deck. Plenty of these calls end with "you don't need a consultant for that".

  2. A scoped proposal

    Within a few days: the approach, the deliverables, the risks that would change the estimate, and a fixed price or day rate. One page, not thirty.

  3. Build, then hand over

    Weekly working software in your repositories, a direct line for questions, and a handover with runbooks, tests and infrastructure as code. Success is your team running it without me.

Engagement models, pricing shapes and what a first month looks like

04 Stack

Tools chosen for the problem, not the CV.

A working knowledge of the modern data stack matters less than knowing which third of it you can safely skip. These are the tools in regular use — grouped the way decisions actually get made.

Processing & streaming
Apache SparkApache Flink KafkaKafka Streams Apache BeamPulsar RisingWaveMaterialize Ray
Storage & query
Apache IcebergDelta Lake Apache HudiSnowflake BigQueryDatabricks RedshiftClickHouse Apache DruidDuckDBTrino
Orchestration & transformation
DagsterApache Airflow PrefectTemporal dbtSQLMesh
Platform & observability
TerraformKubernetes DockerHelmGitOps OpenLineageGreat Expectations DataHubPrometheus GrafanaOpenTelemetry

The full stack, and when each tool is the wrong choice

06 Questions

Before you get in touch

The questions that come up on almost every first call. If yours isn't here, ask it on the call — it's free and it's 30 minutes.

Start a conversation

What size of engagement makes sense?
Anything from a two-week architecture review to a multi-month platform build. The smallest useful engagement is usually a review: a week of reading your pipelines, your warehouse bill and your incident history, ending in a written plan you can execute with or without further help.
Do you work with teams that already have data engineers?
Often. Three shapes recur: a senior pair of hands for a migration nobody has time to run, an architecture reviewer before a large commitment, or someone to build the streaming layer while the in-house team keeps the batch platform running.
Which cloud do you work in?
AWS, Azure and GCP, plus open-source stacks on Kubernetes. The decisions that actually matter are storage layout, table format and orchestration — the cloud is usually whichever one you already pay for.
Can you build things other than data infrastructure?
Yes. Big data is the specialism, but a data platform rarely ships alone. APIs, internal tools, front-ends, cloud infrastructure and CI/CD are part of the same delivery — which means no hand-off between vendors at the seams where projects usually break.
How does pricing work?
Three shapes: a fixed price for a defined build or review, a day rate for open-ended work, or a monthly retainer for ongoing ownership. Every proposal states up front what would change the number, so an estimate doesn't quietly become a surprise.
What does handover look like?
Infrastructure as code, tests, runbooks for the failure modes that actually occur, and a recorded walkthrough. The measure of a good engagement is that your team can run the platform without the consultant who built it.

Next step

Let's talk about your data platform.

Thirty minutes, no charge, no deck. Bring a problem — a slow pipeline, a warehouse bill, a migration that stalled — and you'll leave with a straight opinion on what to do about it.