Scope before you code: what CRISP-DM taught me
Most data projects don’t fail on the model. They fail earlier: on the question asked. Why I spend so much time on scoping.
Abraham Anaba · 4 min read
While taking IBM’s Data Science Methodology course, inspired by John Rollins’s work, one thing struck me: the method devotes its entire first stages to understanding the business problem, long before touching any data. In that sense it complements CRISP-DM, data science’s classic framework.
The available-data trap
The natural reflex on a data project is to start from what you have: “we have these tables, what can we get out of them?”. It is the surest way to build a dashboard nobody opens.
I go the other way: what decision needs to be made, by whom, and how often? Data comes next, in service of that decision.
My method, in four steps
- Understand the decision: who decides what, how often, and with which success metrics.
- Make the data reliable: sources, quality, gaps; one version of the numbers.
- Build with the users: prototype fast, validate with the business, ship to production.
- Measure adoption: track the agreed metrics, and iterate.
What it changes in practice
At Castor Education, this scoping made it possible to cut the number of tracked metrics and turn them into action triggers rather than numbers to comment on. A good metric answers one simple question: what do I do differently if it moves?