Latest Insights
Exploring the frontiers of AI, software architecture, and engineering.
Pareto Analysis: Finding the 20% That Drives the 80%
Five defect types cause 82% of returns. Three features drive 90% of usage. Pareto analysis turns that concentration from folklore into a prioritized, checkable list.
Read Article →Moving Averages vs. Exponential Smoothing: Smoothing Without Self-Deception
Your 7-day average still shows growth two weeks into a decline. Every smoother lags — the question is how much, and whether you chose it deliberately.
Read Article →Standardize or Normalize? Feature Scaling Before You Cluster or Compare
You clustered customers on revenue (thousands) and tenure (years) — and tenure contributed nothing. Unscaled features let the biggest unit win every distance.
Read Article →When to Use a Log Scale (and When It Hides the Story)
The same growth curve looks explosive on a linear axis and boring on a log axis. Both are true. Which one you show depends on the question — here is how to choose.
Read Article →Histograms Lie by Default: Binning Choices That Change the Story
Ten bins show a healthy bell curve. Fifty bins reveal two separate populations. The data never changed — only the binning. How to stop histograms from lying.
Read Article →Duplicate Detection: From Exact Matches to Fuzzy Record Linkage
"Jon Smith" vs "John Smyth" at the same address — same customer or two? Exact matching says two. Your revenue report disagrees. A ladder from dedup to linkage.
Read Article →