Most people discover:
df.drop_duplicates()
But before deleting anything, try:
df.duplicated().sum()
This tells you how many duplicate rows exist.
Want to see them?
df[df.duplicated()]
Want to check duplicates based on specific columns?
df[df.duplicated(subset=["email"])]
And here's a useful one:
df[df.duplicated(subset=["email"], keep=False)]
keep=False marks every occurrence of the duplicate.These commands come in handy when you're trying to understand why duplicates exist before removing them.
#Pandas
@datascience_bds