Session Information
22 SES 10 C, AI and Digitalisation in HE
Paper Session
Contribution
In Italian Universities dropout remains a critical issue with relevant social, economic, and institutional implications. Despite recent improvements across all degree types, the phenomenon continues to be significant (ANVUR, 2023), making the understanding of university dropout rates a major concern for higher education institutions.
Given the multifaceted problems associated with dropout, a key issue lies in identifying effective strategies to understand, monitor and predict this phenomenon at an early stage.
Within this framework, particular attention is devoted to predictive models aimed at detecting and preventing academic failure. Previous research has demonstrated the effectiveness of supervised machine learning approaches for this purpose, particularly when applied to administrative and academic records (Berens et al., 2019; Yağcı, 2022). Moreover, machine learning algorithms have been shown to be powerful predictors of student outcomes, supporting their use as early warning indicators (Delogu et al., 2024), thanks to their ability to capture complex, non-linear relationships among academic, socio-demographic, and other contextual variables. However, predictive accuracy alone is not sufficient when models are intended to support early intervention policies. In this context, ensuring fairness in machine learning models applied to educational data has become relevant (Raftopoulos et al., 2025), especially in settings characterized by heterogeneous populations. Therefore, fairness in early warning systems should focus on ensuring that predictive models do not systematically disadvantage specific groups while preserving data by respecting observed differences across groups. In such settings, calibration metrics are often considered appropriate (Kleinberg et al., 2016).
Given the rapid evolution of these models, other contributions proposed innovative models by applying it to Italian university case (Cannistrà et al., 2022; Ragni et al., 2024; Zanellati et al., 2024). Several studies have also highlighted the role of prior educational pathways and early academic performance (Priulla et al., 2024).
The availability of a nationwide dataset on higher education, including school information and INVALSI large-scale assessments data, represents a valuable resource for better understanding the key factors associated with university path and dropout. After comparing several machine learning models, that have been shown to perform particularly well with this type of data and for classification problems, the aim of this study is to analyze which model could represent a robust baseline for future research on the risk of academic failure. In particular, the study focuses on examining how information available at enrolment and during the first academic year affects model performance, by combining traditional feature importance measures with SHAP (SHapley Additive exPlanations), and on the use of fairness metrics as a diagnostic tool to assess whether boosting models provide equally reliable and actionable risk estimates across groups with different rates of the outcome of interest.
Method
This study is based on a cohort of students derived from a dataset constructed through the integration of multiple data sources including the Italian Ministry of education, National University Register and the National Institute for the Evaluation of the Education System (INVALSI). The cohort, graduated in 2018-2019 school year, is enrolled at university the year after. It was longitudinally followed for five years, allowing the observation of students’ academic progress and graduation outcomes for Italian Bachelor’s degrees collected in the University Register. Different types of information are included such as individual sociodemographic characteristics, upper-secondary school track, INVALSI WLE scores in Italian and Mathematics at grade 13, as well as university-level variables, including the number of credits earned at first academic year, which is a key feature for dropout risk, and changes in field of study. Additionally, student unsuccess is calculated as either failure to graduate on time or withdrawal from university (the latter representing a numerically limited condition in the available data) using information available. A categorical proxy for the distance between the student’s residence and the university is also included in 4 categories. Given that boosting-based models demonstrated superior performance compared to other machine learning approaches, this study focuses on Extreme Gradient Boosting (XGBoost), known for its computational efficiency and predictive accuracy, and CatBoost, which natively and efficiently handles categorical features. These models offer a suitable balance between predictive performance and interpretability, which is particularly relevant in policy-oriented applications. After the preprocessing phase, these models were trained and evaluated to enhance classification performance and improve the overall predictive process. Specifically, a 5-fold cross-validation procedure was adopted, whereby the dataset was split into five folds, with an 80%–20% training–test partition used at each iteration. Hyperparameter tuning was performed using a Random Search strategy (Bergstra & Bengio, 2012), which has been shown to be more efficient approaches in high-dimensional hyperparameter spaces. Model evaluation was conducted not only in terms of overall predictive performance, but also by assessing group-specific metrics in order to examine whether the resulting risk estimates were comparably reliable across student groups (fairness metrics). Finally, model interpretability was addressed by combining model-based feature importance measures with SHAP-based explanations.
Expected Outcomes
The results indicate that model performance is stable across different hyperparameter configurations, suggesting a robust predictive framework in the examined context of the proposed models. XGBoost and CatBoost showed comparable results, reaching approximately 0.77 accuracy and an AUC of 0.83, indicating a robust predictive approach. School-related variables play a relevant role in shaping students’ academic trajectories from enrollment. In particular, final upper-secondary school grade and INVALSI scores are among the top predictors identified through the features importance analysis for all models. Consistent with previous literature, the number of credits earned during the first academic year emerges as a key predictor of students’ academic outcomes. In addition, career events at first academic year (such as changing of faculty, course, etc.) have strong impact on outcome. The integration of SHAP-based method enhances the interpretability of model outputs. Fairness analysis was conducted on model predictions across gender groups. For both XGBoost and CatBoost, no substantial differences were observed in performance metrics, and no evidence of systematic disadvantage in the early warning setting was detected. The fairness analysis highlights the need to complement overall predictive performance with group-specific evaluation metrics, in order to ensure that risk estimates remain reliable across student groups. The findings support the use of machine learning models, particularly boosting-based approaches, as a possible baseline for predicting student career, providing a neutral setting that can be adopted in supporting first-year intervention strategies, while emphasizing the need for careful and responsible use of predictive tools. The integration of multiple administrative data sources at the national level enhances the ability to analyse key factors related to university academic trajectories.
References
ALMALAUREA (2024). Rapporto 2024 sul profilo e sulla condizione occupazionale dei laureati, https://www.almalaurea.it ANVUR. (2023). Rapporto biennale sullo stato del sistema universitario e della ricerca. https://www.anvur.it/it/dati-e-pubblicazioni/rapporto-biennale. Atzeni, G., et al. (2022). Drop-Out Decisions in a Cohort of Italian Universities. In: Teaching, Research and Academic Careers: An Analysis of the Interrelations and Impacts. Cham: Springer International Publishing, pp. 71-103. Berens, J., Schneider, K., Gortz, S., Oster, S., & Burghoff, J. (2019). Early Detection of Students at Risk--Predicting Student Dropouts Using Administrative Student Data from German Universities and Machine Learning Methods. Journal of Educational Data Mining, 11(3), 1-41. Bergstra, J., & Bengio, Y. (2012). Random search for hyper-parameter optimization. The journal of machine learning research, 13(1), 281-305. Campodifiori, E., Figura, E., Papini, M., & Ricci, R. (2010). Un indicatore di status socioeconomico-culturale degli allievi della quinta primaria in Italia (Working Paper No. 2). INVALSI. http://www.invalsi.it/download/wp/wp02_Ricci.pdf. Cannistrà, M., Masci, C., Ieva, F., Agasisti, T., & Paganoni, A. M. (2022). Early-predicting dropout of university students: an application of innovative multilevel machine learning and statistical techniques. Studies in Higher Education, 47(9), 1935-1956. Delogu, M., Lagravinese, R., Paolini, D., & Resce, G. (2024). Predicting dropout from higher education: Evidence from Italy. Economic Modelling, 130, 106583. Hastie, T., Tibshirani, R., & Friedman, J. H. (2013). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer: Berlin/Heidelberg, Germany. Kleinberg, J., Mullainathan, S., & Raghavan, M. (2016). Inherent trade-offs in the fair determination of risk scores. arXiv preprint arXiv:1609.05807. Kuhn, M., & Johnson, K. (2013). Applied Predictive Modeling. Springer. Priulla,A., Albano, A., D'Angelo, N., & Attanasio, M. (2024). A machine learning approach to predict university enrolment choices through students' high school background in Italy. ArXiv abs/2403.13819. Ragni, A., Masci, C., & Paganoni, A. M. (2024). Analysis of Higher Education Dropouts Dynamics through Multilevel Functional Decomposition of Recurrent Events in Counting Processes. arXiv preprint arXiv:2411.13370. Raftopoulos, G., Davrazos, G., & Kotsiantis, S. (2025). Evaluating fairness strategies in educational data mining: A comparative study of bias mitigation techniques. Electronics, 14(9), 1856. Yağcı, M. Educational data mining: prediction of students' academic performance using machine learning algorithms. Smart Learn. Environ. 9, 11 (2022). Zanellati, A., Zingaro, S. P., & Gabbrielli, M. (2024). Balancing performance and explainability in academic dropout prediction. IEEE Transactions on Learning Technologies, 17, 2086-2099.
Update Modus of this Database
The current conference programme can be browsed in the conference management system (conftool) and, closer to the conference, in the conference app.
This database will be updated with the conference data after ECER.
Search the ECER Programme
- Search for keywords and phrases in "Text Search"
- Restrict in which part of the abstracts to search in "Where to search"
- Search for authors and in the respective field.
- For planning your conference attendance, please use the conference app, which will be issued some weeks before the conference and the conference agenda provided in conftool.
- If you are a session chair, best look up your chairing duties in the conference system (Conftool) or the app.