Blog

Latest Insights

Exploring the frontiers of AI, software architecture, and engineering.

Sep 22, 2026

Pareto Analysis: Finding the 20% That Drives the 80%

Five defect types cause 82% of returns. Three features drive 90% of usage. Pareto analysis turns that concentration from folklore into a prioritized, checkable list.

Read Article →
Sep 20, 2026

Moving Averages vs. Exponential Smoothing: Smoothing Without Self-Deception

Your 7-day average still shows growth two weeks into a decline. Every smoother lags — the question is how much, and whether you chose it deliberately.

Read Article →
Sep 18, 2026

Standardize or Normalize? Feature Scaling Before You Cluster or Compare

You clustered customers on revenue (thousands) and tenure (years) — and tenure contributed nothing. Unscaled features let the biggest unit win every distance.

Read Article →
Sep 16, 2026

When to Use a Log Scale (and When It Hides the Story)

The same growth curve looks explosive on a linear axis and boring on a log axis. Both are true. Which one you show depends on the question — here is how to choose.

Read Article →
Sep 14, 2026

Histograms Lie by Default: Binning Choices That Change the Story

Ten bins show a healthy bell curve. Fifty bins reveal two separate populations. The data never changed — only the binning. How to stop histograms from lying.

Read Article →
Sep 12, 2026

Duplicate Detection: From Exact Matches to Fuzzy Record Linkage

"Jon Smith" vs "John Smyth" at the same address — same customer or two? Exact matching says two. Your revenue report disagrees. A ladder from dedup to linkage.

Read Article →