Raw software inventory is messy. The same product shows up under a dozen names, spelled differently by every source system. One device reports "Microsoft Office Professional Plus 2019", another "MS Office Pro Plus 2019 (x64)", a purchase record says "Office 2019 ProPlus", and a SaaS invoice simply says "Microsoft". Until those records are recognized as the same thing, every count, report and license position built on them is wrong.
Software normalization is the work of turning that mess into a clean, consistent list of publishers, products, editions and versions.
Why raw data is so messy
Every source describes software in its own way:
- Device inventory reports what the installer registered, including updates, language packs and components.
- SaaS admin consoles describe subscriptions by plan or SKU name.
- Purchase and invoice data use the reseller's or vendor's billing descriptions.
- Contracts use the publisher's legal product names, which change with rebranding.
Mergers and rebranding add to the confusion. Products change owners and names, and old records keep the old ones.
What normalization does
A normalization process works in layers:
- Remove noise
Exclude updates, hotfixes, drivers and components that are not separately licensed.
- Merge duplicates
Recognize the same installation reported by more than one source.
- Normalize the publisher
"MSFT", "Microsoft Corp" and "Microsoft Corporation" become Microsoft.
- Normalize the product and edition
Map variants to one product name and the edition that determines the license.
- Group versions where it matters
Keep version detail only where license rights depend on it.
- Flag what is licensable
Separate free and open-source software from products that need entitlements.
Rules that keep normalization consistent
- Maintain one master catalog of publisher and product names, and map everything to it.
- Keep the raw name alongside the normalized one, so mappings can be audited and corrected.
- Normalize once, at intake, rather than in each report.
- Review unmatched names every month and add rules for them.
- Map rebranded and acquired products to their new name.
- Fixing names by hand inside individual reports.
- Overwriting the raw name, so nobody can check the mapping later.
- Creating a second product when a product is renamed.
- Normalizing every version of everything, including free tools nobody licenses.
Where normalization pays off
Clean names are not an end in themselves. They make three things possible:
Entitlements and usage can only be compared when both refer to the same product. See the effective license position.
Finance can see total spend per publisher across contracts, invoices and expense claims. See spend analysis.
When ten teams use ten tools for the same job, normalized categories make the overlap visible. See application rationalization.
How MI One helps
Frequently asked questions
Is normalization the same as software recognition?
They are closely related. Recognition identifies what a raw record is; normalization maps it to a standard name and structure. Most tools do both.
Do we need to normalize every version?
No. Keep version detail only where license rights depend on it, for example when a license covers one major version and not the next.
How long does normalization take?
For a mid-sized estate, an initial pass on the top vendors takes days rather than weeks. Keeping it current takes a short monthly review.