Studying at the University of Verona
Here you can find information on the organisational aspects of the Programme, lecture timetables, learning activities and useful contact details for your time at the University, from enrolment to graduation.
Study Plan
The Study Plan includes all modules, teaching and learning activities that each student will need to undertake during their time at the University.
Please select your Study Plan based on your enrollment year.
1° Year
| Modules | Credits | TAF | SSD |
|---|
2° Year It will be activated in the A.Y. 2026/2027
| Modules | Credits | TAF | SSD |
|---|
| Modules | Credits | TAF | SSD |
|---|
| Modules | Credits | TAF | SSD |
|---|
| Modules | Credits | TAF | SSD |
|---|
1 module among the following2 modules among the following1 module among the following
A.A. 2025/26 e 2026/27 - Network science and econophysics not activated
A.A. 2026/27 Complex systems and social physics not activated1 module among the following2 modules among the followingLegend | Type of training activity (TTA)
TAF (Type of Educational Activity) All courses and activities are classified into different types of educational activities, indicated by a letter.
Statistical models for Data Science (2025/2026)
Teaching code
4S009079
Academic staff
Coordinator
Credits
6
Also offered in courses:
- Statistical Models of the course Master's degree in Artificial intelligence
- Statistical Models of the course Master's degree in Artificial intelligence
- Statistical models for Data Science of the course Master's degree in Mathematics
Language
English
Scientific Disciplinary Sector (SSD)
MAT/06 - PROBABILITY AND STATISTICS
Period
1st semester dal Oct 1, 2025 al Jan 30, 2026.
Courses Single
Authorized
Learning objectives
The course will be devoted to the mathematical background necessary to describe, analyze and derive value from datasets, possibly Big Data and unstructured, and to master the main probabilistic models used in the data science field. Starting from basic models, for example regressions, PCA-based predictors, Bayesian statistics, filters, etc., particular emphasis will be placed on mathematically rigorous quantitative approaches aimed at optimizing the data collection, cleaning and organization phases (e.g. series historical data, unstructured data generated in social media, semantic elements, etc.). The mathematical tools necessary to deal with the description of the time series, their analysis and forecasts will also be introduced. The contents of the entire course will be structured in interaction with the study of real problems relating to industrial, economic, social, etc., heterogeneous sectors, using software oriented to probabilistic modeling, for example, Knime, ElasticSearch, Kibana, R AnalyticFlow, Orange , etc.
Prerequisites and basic notions
Relative to both modules that make up the entire course:
Basic notions of Probability theory, e.g., definition of probability space, definition of random variables, even in multiple dimensions, concept of independence between random variables
Knowledge of the main models of notable random variables, both discrete and continuous, e.g., binomial, Poisson, Gaussian, and their main statistical properties
Convergence theorems, e.g., the law of large numbers, the central limit theorem
Basic notions of discrete-time and continuous-time stochastic processes, e.g., Markov chains, birth and death processes Rudiments of statistical and data analysis, e.g., concepts of frequency, mean, mode, and squared deviation
Basic notions of programming in Python, e.g., general syntax, data structures, import/export, and construction of graphs for data visualization
Rudiments of the main Python-related libraries, e.g., Numpy, Pandas, and Matplotlib.
Program
The course program is divided into the following macro-topics. Part 1 [ module 1 ]
1. Time domain analysis
2. Frequency domain analysis
3. Tools for data analysis and cleaning, e.g. outlier identification 4. Maximum likelihood methods, likelihood metrics, probability density fitting and Principal Component Analysis (PCA)
5. AR, MA, ARMA, ARIMA, Box-Jenkins, ARCH, GARCH models and generalizations
6. Time decomposition tools for chronologically ordered data series
7. Hypothesis testing
8. Gaussian / jump/compound stochastic processes
9. Decomposition of stochastic processes of the "white noise" type
10. Bayesian statistics and applications
11. Forecasting using statistical inferential models, based on, e.g., the concepts of autocovariance and partial autocorrelation, seasonality (SARIMA), analysis of variance (ANOVA, MANOVA), etc.
12. Smoothing techniques, spectral decomposition, polynomial fitting,
Part 2 [module 2]
1. Python programming review
2. Managing and visualizing time series
3. Descriptive statistics
4. Frequency domain analysis
5. Linear regression for time series
6. Analyze and decompose the principal components of time series (trend, cycle, seasonality)
7. Forecasting methods: Exponential Smoothing (simple, double, triple)
8. Forecasting methods: AR, MA, ARMA, ARIMA, SARIMA
9. Forecasting methods: ARCH, GARCH, and generalizations
10. How to evaluate different forecasting models.
All the above points will be explored in depth through practical exercises that require the implementation of appropriate Python code. Furthermore, the main forecasting methods will be further explored thanks to the treatment and resolution of real case studies of various types.
Bibliography
Didactic methods
The course will be divided into lectures, with slides as well as notes sharing, and computer simulations / exercises.
Learning assessment procedures
The final exam consists of two parts: one theoretical, the next practical / implementative. Consequently, the first part of the exam is functional to the verification of the learning of the theoretical concepts characterizing the statistical methods and the connected models and algorithms, at the basis of the IT-computational implementations used to donduct a project that the student will agree with the course teachers.
Latter "case study", together with the discussion of the coding parts created to complete it, will be the subject of the second and final part of the exam.
Evaluation criteria
The evaluation of the exam will be carried out by combining the results obtained from the two modules of the course, therefore giving equal importance to the correctness and effectiveness of the solutions adopted in the phase of solving concrete problems due to computer implementations, as well as to understanding of the probabilistic / statistical models underlying them.
Criteria for the composition of the final grade
The final grade will be the result of a joint evaluation of the two tests: the theoretical one and the second one focused on the resolution of a "case study", agreed upon by the student with the teachers and in accordance with what is expressed in the "Exam Method", and "Assessment Criteria" sections.
Exam language
Inglese / English
