How raw data was cleaned, transformed, and prepared for machine learning.
Exact duplicate rows can inflate cluster sizes artificially.
We use drop_duplicates() to ensure each customer record is unique.
CustomerID is just an identifier — it carries no business information and would confuse the clustering algorithm if included.
Renaming columns to Python-friendly names (no spaces, no special characters) makes code cleaner and avoids indexing issues.
Machine learning algorithms require numeric inputs. We encode Gender as a binary column: Male=1, Female=0. This is standard label encoding for binary categoricals.
After all cleaning steps, we verify: no remaining nulls, correct dtypes, and correct shape.
| Gender | Age | Annual_Income | Spending_Score | Gender_Encoded |
|---|---|---|---|---|
| Male | 19 | 15 | 39 | 1 |
| Male | 21 | 15 | 81 | 1 |
| Female | 20 | 16 | 6 | 0 |
| Female | 23 | 16 | 77 | 0 |
| Female | 31 | 17 | 40 | 0 |
| Female | 22 | 17 | 76 | 0 |
| Female | 35 | 18 | 6 | 0 |
| Female | 23 | 18 | 94 | 0 |
| Male | 64 | 19 | 3 | 1 |
| Female | 30 | 19 | 72 | 0 |