Insights, tutorials, and thoughts on technology
A working reference for the commands you actually reach for day to day — cluster inspection, table design, loading, and troubleshooting. Copy, adapt, run.
The commands and `bq` CLI incantations that come up constantly once you're running BigQuery in production — not the beginner tutorial version.
The SQL and SnowSQL commands that come up constantly once you're running a real Snowflake environment — warehouses, RBAC, loading, and monitoring.
The Spark SQL, Delta Lake, and PySpark commands that come up constantly in day-to-day Databricks work.
Most Redshift performance problems trace back to one of four things: bad distribution, unsorted data, oversized WLM queues, or queries fighting each other for t
Snowflake's separation of storage and compute makes it forgiving in ways Redshift and Databricks aren't — but that same flexibility means it's easy to paper ove
Databricks performance problems tend to cluster around three sources: small files, poor partitioning/Z-ordering, and cluster configurations that don't match the
BigQuery's serverless model removes cluster management, but it replaces "is my cluster big enough" with a different question: "how much data is this query actua
Teradata and SQL Server served enterprises well for decades, but both were built for a world of predictable batch windows and structured-only data. Databricks is built for a different world.
Every Teradata-to-Redshift migration deck looks the same: extract, land in S3, load with COPY, done. The diagram isn't wrong, it's just missing the ten decisions that actually determine success.
Snowflake ships new features constantly, and most write-ups just list them. The more useful question for a data engineering team is: which are worth your implementation time this quarter?
Cloud-to-cloud data warehouse migrations get harder than same-cloud migrations for a reason people underestimate: you're not just changing SQL dialects, you're changing the entire operating model.