Big data engineering · Contract & freelance · Remote, EU
Data platforms that survive contact with production.
Quarkray is a one-person engineering practice built around large-scale data. Pipelines, lakehouses and streaming systems — scoped, built and handed over by the same senior engineer. No account managers, no junior hand-off, no six-week discovery phase before anyone writes code.
Straight answer on the call, written scope within a few days.
The engineer who scopes the work is the one who writes it, and the one who hands it over.
Big data platforms as the specialism; APIs, apps and cloud infrastructure when the platform needs them.
Fixed-scope build, architecture review, or monthly retainer. Contract or freelance, remote.
Runbooks, tests, infrastructure as code and a walkthrough — so your team can own it.
01 The usual symptoms
Most data problems don't start as data problems.
They start as a deadline. Something ships that works on Tuesday's data. Then the volume doubles, an upstream schema changes without warning, and the pipeline becomes the thing nobody wants to touch. These four show up in almost every audit.
The pipeline that fails quietly
A job dies at 3am and a stakeholder finds out first. The missing piece is rarely monitoring — it's idempotent writes, real retries, and a data contract with the upstream team so a renamed column stops being an outage.
The bill that outgrows the data
Warehouse spend climbing while query patterns stay flat. It's usually file sizes, partitioning and clustering — plus a handful of dashboards quietly scanning full tables on a five-minute refresh.
Two numbers for one metric
Finance and product disagree because "active customer" is defined in five dashboards and two notebooks. That's a modelling problem with a known fix: one tested definition, in version control, with lineage you can point at.
The migration stuck at 80%
A lakehouse or warehouse move that has been nearly done for two quarters. The target architecture was designed; the cut-over — dual writes, reconciliation, the order things get switched off — never was.
02 Services
Deep in data. Broad enough to ship the whole thing.
The specialism is big data engineering — platforms that handle real volume without a team of five to babysit them. The range covers everything that platform has to connect to, so there's no hand-off between vendors at exactly the seams where projects break.
Data platform architecture
A written architecture you can build against: ingestion, storage layout, modelling, orchestration and access — sized to your team and budget, not to a vendor reference diagram.
Read more → S/02Pipelines & orchestration
Batch and change-data-capture pipelines that are idempotent, backfillable and observable. Airflow, Dagster or Prefect — picked for your operational reality, not for fashion.
Read more → S/03Real-time & streaming
Kafka, Flink and streaming SQL for the cases that genuinely need seconds instead of hours: fraud checks, live pricing, operational dashboards, event-driven services.
Read more → S/04Lakehouse & warehouse
Iceberg, Delta or Hudi on object storage; Snowflake, BigQuery or Databricks where they earn their cost. Migrations planned as cut-overs with reconciliation, not as big bangs.
Read more → S/05Analytics engineering
dbt or SQLMesh models with tests, lineage and exactly one definition per metric — so the number in the board deck matches the number in the product.
Read more → S/06AI & LLM data infrastructure
The unglamorous nine tenths of an AI feature: chunking, embeddings, vector storage, retrieval evaluation, and the pipelines that keep the index fresh as the source data moves.
Read more →Also built when the platform needs it: FastAPI and Node services, React and Next.js front-ends, Terraform, Kubernetes and CI/CD. Full service detail →
03 How it works
Three steps, no discovery theatre.
-
A 30-minute call
You describe the problem. You get a straight opinion on whether it's worth building, buying, or leaving alone — free, and without a slide deck. Plenty of these calls end with "you don't need a consultant for that".
-
A scoped proposal
Within a few days: the approach, the deliverables, the risks that would change the estimate, and a fixed price or day rate. One page, not thirty.
-
Build, then hand over
Weekly working software in your repositories, a direct line for questions, and a handover with runbooks, tests and infrastructure as code. Success is your team running it without me.
Engagement models, pricing shapes and what a first month looks like →
04 Stack
Tools chosen for the problem, not the CV.
A working knowledge of the modern data stack matters less than knowing which third of it you can safely skip. These are the tools in regular use — grouped the way decisions actually get made.
05 Writing
Notes from production.
Practical write-ups on the problems that actually consume a data team's week: tuning, architecture trade-offs, and the parts of AI infrastructure nobody demos.
10 Apache Spark performance tuning tips
Partition sizing, Adaptive Query Execution, shuffle costs and the caching decisions that quietly make jobs slower.
Real-time streaming analytics with AI: a practical guide
Building streaming architectures with Kafka and Flink, and where real-time inference belongs in them.
How big data pipelines power modern AI systems
Data lakes, feature stores and training pipelines — the backbone under every production model.
06 Questions
Before you get in touch
The questions that come up on almost every first call. If yours isn't here, ask it on the call — it's free and it's 30 minutes.
What size of engagement makes sense?
Do you work with teams that already have data engineers?
Which cloud do you work in?
Can you build things other than data infrastructure?
How does pricing work?
What does handover look like?
Next step
Let's talk about your data platform.
Thirty minutes, no charge, no deck. Bring a problem — a slow pipeline, a warehouse bill, a migration that stalled — and you'll leave with a straight opinion on what to do about it.