Listen to the audio version on Spotify and Apple Podcasts.
In our last edition, we looked at the big picture: what useful clusters look like in the context of disability innovation research. Now, let's look at how to actually build them—which means stepping into the practical, everyday world of Microsoft Excel. We are using Excel, but our approach can definitely be applied to other spreadsheets or even databases.
Within the Zero Project team, we have several Excel aficionados with strong preferences of how to organize those tables that we use for collaborations. Trust me, debates over hidden cells, dropdown menus, and column layouts can get surprisingly heated! (Out of years of manually managing large, diverse datasets, we have set a few foundational rules that keep our data clean and our team sane.
Rows vs. columns
Once you have made your decisions on your clusters, there is an easy way to implement, adapt, and develop them. It is called an Excel column, and this column marks those projects (or persons) with “yes” that are in the cluster.
Usually, you put your projects in the rows and the clusters – next to all other categorizations, descriptions, codes, etc. – in the columns.
When it makes sense to turn your spreadsheet sideways
Your spreadsheets get unwieldy when you want to add a long description of a project into one column of your spreadsheet, and this one column contains many more characters than all the others. (In this case, your spreadsheet stays more manageable if the projects are in the columns and their categories are in the rows.
We know that this is against standard rules of most spreadsheet artisans, but in practice this has proven to work: Our MASTERFACTSHEET that contains the final version of all Factsheets, has the projects in the columns).
From individual clusters to combinations
Try to keep all your clusters separate and unrelated to each other. Add, for example, one column for "Region," which then allows for the following clusters: Europe, Asia, Africa, America, and Oceania. Then, add another column if you want to subdivide the world into smaller regions.
Always use dropdown fields and only work with predefined fields. This saves you a lot of double-checking and protects you from those team members who can never agree whether to write USA, U.S., United States, or United States of America – you define the only version that the spreadsheet accepts.
Combining clusters for richer insights
You may want to add another column for country income, for example, according to the UNDP: very high, high, middle, low. And another column for OECD member (yes/no), or another for EU member (yes/no). This allows you to filter and disaggregate your table into any cluster and cluster combination you want, like “Asian countries with high or very high income”.
Is it education, or technology, or both?
The “separate-column approach” also takes care of the fact that your clusters do not usually slice the whole picture into nicely fitting puzzle pieces. Usually, clusters overlap, and they also often leave gaps in between. Some clusters will be fully or partly contained in another. Autism will be part of neurodiversity, and you should still use both, since you might need both clusters for different purposes.
In general, the world of disability is a world with fuzzy definitions. Neurodiversity, for example, has no commonly accepted list of diagnoses that the term covers.
Never force a project into a single category
Most importantly, never use a clustering system that can only work with a singular categorization system. Imagine when, during clustering, you have to ask yourself the question: "Is this project (more) education or (more) assistive technology? And I am in a dilemma now, because it has elements of both." If that happens, you need to re-start and find a different approach to your clustering, usually using one clustering column for education and the other for AT. Now one project can be both education and AT.
Forcing singular decisions on your data always ends up with somewhat arbitrary, inconsistent, incomplete, and competing clusters. One day, you will tend more towards education, and the next towards AT. If you see any signs of that: change your clustering system!
Empty clusters?
Sometimes an interesting question will come up when you have defined a cluster that makes sense to you but stays empty because none of your projects belong to it. So, should you stick with clusters that are empty (or have only one project in them, which is also odd for a cluster)? The opposite may also happen: should you stick with clusters that contain 80 percent or more of all projects?
When cluster sizes tell a story
For some purposes, it is useless or at least detrimental to have clusters that are (almost) empty, or that contain a vast majority of projects. So, you should be on the lookout for more useful clusters, such as subdividing large ones or merging small ones. Or it might also mean that your clusters were misconceptions from the beginning and that the real world of existing projects is different from what you expected it to be.
At the Zero Project, we used a technology segmentation that we borrowed from a European Union paper that we considered the gold standard at that time (10 years ago). It contained clusters like Internet of Things or Big Data.
When outdated clusters need to evolve
Not only did we always receive a small number of projects that fit into these clusters, but technology has moved on and, in the context of accessibility and disability inclusion, clearly other clusters are much more useful.
But there might also be reasons to stick to such highly (or lightly) populated clusters, depending on the use case. It may, for example, give a clear indication that within certain clusters a lot of innovation is happening, whereas in others not so much. . That is very useful information, which could be watered down by changing the clusters to even out their sizes.
AI clustering power at your fingertips
Looking into the near future, working with clusters is currently being revolutionized by ChatGPT, Claude, and the like.
Although I studied statistics at university and applied some of the methods that I learned all the time, like probabilities and correlations, I never used the statistical method called “cluster analysis”. Simply because it was too complicated – until now: With new AI tools, auto-clustering and auto-detection of clusters are at your fingertips now. I have not worked it out systematically yet but will come back to it in a future blog, when it has proven useful for our practical work at the Zero Project.
What are your experiences with AI tools for clustering datasets?
read more