Fall 2026
Data Mining
This course covers the analysis and extraction of patterns in datasets using unsupervised machine learning methods and advanced techniques. Students will learn to structure analytical projects and apply ensemble methods, clustering, association rules, and dimensionality reduction. The approach combines theoretical foundations with hands-on implementation in Python to solve data problems in scientific and technological domains.
Staff

Syllabus
Download the full syllabus as a PDF.
Syllabus (PDF)Topics
Introduction and Data Preparation
- Introduction to Data Mining
- Exploratory Data Analysis: Cleaning, Imputation, Scaling
- Outlier Detection (Z-Score, IQR)
Ensemble Methods
- Decision Trees and Gini Impurity
- Ensembling via Bagging, Boosting, and Random Forest
- Feature Subspaces and Error Metrics
Clustering
- Clustering Concepts and Distances
- K-Means Algorithm and k-means++ Initialization
- Cluster Evaluation (Inertia and Silhouette)
- Density-Based Clustering: DBSCAN Algorithm
- Selecting Epsilon via k-distances and Noise
Association Rules
- Transactional Analysis and the Apriori Algorithm
- Rule Metrics: Support, Confidence, and Lift
- Multi-Objective Filtering and the Pareto Frontier
Dimensionality Reduction
- Covariance Matrix, Eigenvalues, and Eigenvectors
- Orthogonal PCA Projection and Explained Variance
- Visualization in Non-Linear Spaces (t-SNE and UMAP)
Grading
- 80%
Programming Assignments
6 hands-on Python challenges developed throughout the semester (every 2–3 weeks).
- 20%
Final Exam
In-person written exam at the end of the semester, consisting of technical-interview-style exercises focused on problem solving, pseudocode, algorithmic analysis, and theoretical foundations.
Policies
- Attendance carries no direct percentage weight in the final grade. However, a minimum of 80% attendance is required to be eligible for continuous assessment and the final ordinary exam.
- All submitted work (assignments, presentations, and projects) must be original. Generative AI tools may be used only for research, brainstorming, or grammar review, but not to draft the final content. The instructor reserves the right to request an in-person oral defense of any submitted work; if the student cannot demonstrate mastery of the topic or authorship of the text during that defense, the activity will be voided (a grade of zero).
Readings
Available at the UNAM digital library (bidi.unam.mx) as downloadable PDFs.
- Aggarwal, C. C. (2015). Data Mining: The Textbook. Springer.
- Aggarwal, C. C., & Reddy, C. K. (2013). Data Clustering: Algorithms and Applications. CRC Press / Chapman & Hall.
- Bramer, M. (2020). Principles of Data Mining (4th ed.). Springer.
- Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
- Tan, P.-N., Steinbach, M., Karpatne, A., & Kumar, V. (2018). Introduction to Data Mining (2nd ed.). Pearson.
- James, G., Witten, D., Hastie, T., Tibshirani, R., & Taylor, J. (2023). An Introduction to Statistical Learning: with Applications in Python (ISLP). Springer.
- Géron, A. (2022). Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly Media.
- McKinney, W. (2022). Python for Data Analysis: Data Wrangling with Pandas, NumPy, and Jupyter (3rd ed.). O'Reilly Media.
- VanderPlas, J. (2022). Python Data Science Handbook: Essential Tools for Working with Data (2nd ed.). O'Reilly Media.
- Zaki, M. J., & Meira Jr, W. (2020). Data Mining and Machine Learning: Fundamental Concepts and Algorithms (2nd ed.). Cambridge University Press.
- Deisenroth, M. P., Faisal, A. A., & Ong, C. S. (2020). Mathematics for Machine Learning. Cambridge University Press.
