Notes from production
Practical write-ups on the problems that actually consume a data team's week — tuning, architecture trade-offs, streaming, and the parts of AI infrastructure nobody demos. No listicles about the future of data.
05 Articles
10 Apache Spark performance tuning tips
Partition sizing, Adaptive Query Execution, shuffle cost and caching — ten levers that decide whether a Spark job takes minutes or hours.
Real-time streaming analytics with AI: a practical guide
A hands-on guide to streaming architecture: Kafka ingestion, Flink processing, state and windowing, and where real-time model inference actually belongs.
How big data pipelines power modern AI systems
Data lakes, feature stores and training pipelines are the backbone under every production model. What the pipeline layer of an AI system has to get right.
AI agents: the next frontier in enterprise automation
Agentic architectures, tool use and multi-agent systems — what actually changes for enterprise automation, and what still needs a boring pipeline underneath.
The essential AI skills every data professional needs
Prompt engineering, retrieval, vector databases, MLOps and evaluation — the skills that changed what a data engineer is expected to ship.
What gets written about here
Four threads, chosen because they are where engagements usually start: performance work on existing platforms, streaming architecture, the data layer under AI features, and the operational practices that keep a platform boring.
- Performance & cost — Spark tuning, warehouse spend, file layout
- Streaming — Kafka, Flink, exactly-once, state and windowing
- Architecture — lakehouse table formats, modelling, migrations
- AI infrastructure — retrieval, embeddings, evaluation, freshness
Next step
Got one of these problems for real?
Thirty minutes, no charge, no deck — and a straight opinion on what to do about it.