▎Common Pandas Terms
1. Series: A one-dimensional labeled array capable of holding any data type (integers, strings, floating point numbers, Python objects, etc.).
2. DataFrame: A two-dimensional, size-mutable, and potentially heterogeneous tabular data structure with labeled axes (rows and columns).
3. Index: The labels for the rows of a Series or DataFrame, used for fast identification and alignment of data.
4. read_csv: A widely used function to load data from a Comma-Separated Values file into a Pandas DataFrame.
5. head() / tail(): Methods used to quickly inspect the first or last few rows (default is 5) of a DataFrame or Series.
6. loc: A label-based data selection method used to access a group of rows and columns by their labels or a boolean array.
7. iloc: An integer-location based selection method used to access data by its numerical position (0-based indexing).
8. Shape: An attribute that returns a tuple representing the dimensionality of the DataFrame (number of rows, number of columns).
9. Describe: A method that generates descriptive statistics (mean, count, std, min, max, etc.) for numerical columns in a DataFrame.
10. GroupBy: A process involving splitting the data into groups based on some criteria, applying a function, and combining the results.
11. Aggregation (agg): The process of computing a summary statistic (like sum, mean, or count) for each group in a dataset.
12. Merge: A function used to combine two DataFrames based on a common key or index, similar to a SQL JOIN operation.
13. Concatenation (concat): The process of "gluing" together multiple DataFrames or Series along a particular axis (either rows or columns).
14. dropna: A method used to remove missing values (NaN) from a Series or DataFrame.
15. fillna: A method used to replace missing values (NaN) with a specified value or a calculated value (like the mean or median).
16. Apply: A powerful method that allows you to apply a function along an axis of the DataFrame or on a Series.
17. Pivot Table: A method used to summarize and reshape data into a spreadsheet-style table, often used for multi-dimensional analysis.
18. Melt: A function used to transform a "wide" DataFrame into a "long" format, unpivoting columns into rows.
19. Vectorization: The process of performing operations on entire arrays (columns) at once without the need for explicit Python loops, ensuring high performance.
20. DatetimeIndex: A specialized type of index in Pandas that handles date and time information, enabling powerful time-series analysis and resampling.
Post #1210
606
- ❤ 4