Answered Questions
It depends on the target population. If you include all hospitalized patients during a defined period and your conclusions are limited to those patients, statistical…
It depends on the target population. If you include all hospitalized patients during a defined period and your conclusions are limited to those patients, statistical inference is generally not necessary since you have the whole population of interest. However, if you want to generalize beyond them (e.g., to future patients, patients in other periods, or similar hospitals), statistical tests and confidence intervals can still be appropriate.
Type I error (α): Rejecting a true null hypothesis → false positive.Public health example: Concluding that a vaccination campaign reduces infection rates when it actually…
- Type I error (α): Rejecting a true null hypothesis → false positive.
Public health example: Concluding that a vaccination campaign reduces infection rates when it actually does not. - Type II error (β): Failing to reject a false null hypothesis → false negative.
Public health example: Concluding that a vaccination campaign does not reduce infection rates when it actually does. - Power = 1 − β: The probability of correctly detecting a true effect.
Regression is a statistical method used to examine the relationship between an outcome variable (dependent variable) and one or more predictor variables (independent variables). Regression…
- Linear regression = continuous outcomes
- Logistic regression = binary outcomes
- Poisson regression = count data
- Cox regression = survival/time-to-event outcomes
- Multilevel regression = clustered/hierarchical data
- Fits a statistical relationship between predictors and outcome
- Estimates coefficients showing the direction and strength of associations
- Tests statistical significance
- Evaluates model fit and assumptions
- https://stats.oarc.ucla.edu/r/seminars/introduction-to-regression-in-r/
- https://stats.oarc.ucla.edu/stata/webbooks/reg/chapter1/regressionwith-statachapter-1-simple-and-multiple-regression/
- https://stats.oarc.ucla.edu/spss/seminars/introduction-to-regression-with-spss/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC2992018/
After running regression analysis, we usually check whether the model assumptions are satisfied and whether the model fits the data well. Common checks include: residual…
- Linear regression = continuous outcome
- Logistic regression = binary outcome
- Survival analysis = time-to-event outcome
- Multilevel models = clustered/hierarchical data
- Bayesian models = when incorporating prior information or handling complex uncertainty
Below are some excellent free resources that can help you learn sample size determination and improve your methodological thinking when designing the methodology section of…
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10000262/
- https://iris.who.int/items/9c2e5da4-3785-4fec-9dbc-841e4ae0d98c
- https://www.sciencedirect.com/science/article/pii/S2772906024005089
- OpenEpi Sample Size Calculator
- Methodological Thinking: Basic Principles of Social Research Design: https://methods.sagepub.com/book/mono/methodological-thinking-2e/toc#
- https://www.nature.com/articles/6400355
- https://pubmed.ncbi.nlm.nih.gov/16355244/
- https://www.nature.com/articles/6400375
- https://pubmed.ncbi.nlm.nih.gov/16858385/
- https://pubmed.ncbi.nlm.nih.gov/17003803/
- https://pubmed.ncbi.nlm.nih.gov/17187048/
Comparing time series forecasting models such as ARIMA, SARIMA, LSTM, and Prophet involves training each model on historical data and evaluating their forecasting performance on…
Comparing time series forecasting models such as ARIMA, SARIMA, LSTM, and Prophet involves training each model on historical data and evaluating their forecasting performance on test data. This is done by using metrics such as Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Absolute Percentage Error (MAPE), and Mean Squared Error (MSE). Lower values generally indicate better predictive performance. Briefly, each model works best in different situations.
- ARIMA is best suited for non-seasonal linear trends.
- SARIMA extends ARIMA by handling seasonal patterns.
- Prophet works best for business and public health data with seasonality, trends, and missing observations.
- LSTM (Long Short-Term Memory) is a deep learning model that can capture complex nonlinear temporal relationships, especially in large datasets
- https://otexts.com/fpp3/
- https://www.sciencedirect.com/science/article/pii/S0169207021001874
- https://www.stata.com/features/time-series/
Survival analysis refers to a set of statistical methods used to analyse the time until an event occurs. The “event” includes death, relapse, progression, or…
Survival analysis refers to a set of statistical methods used to analyse the time until an event occurs. The “event” includes death, relapse, progression, or treatment. Different statistical methods are used to answer different questions. The main methods include the following:
- Kaplan-Meier analysis – estimate survival over time
- Log-rank test – compare survival curves between groups
- Cox proportional hazards model – assess effect of predictors on survival while adjusting for confounders.
- Competing risks methods – used when more than one type of event can occur
- Parametric survival models – model survival with assumed distribution
In paediatric cancer research, survival analysis is commonly used to study time until an event such as death, relapse, progression, or treatment failure occurs. Different…
In paediatric cancer research, survival analysis is commonly used to study time until an event such as death, relapse, progression, or treatment failure occurs. Different methods are used to answer different questions. The main methods used include the following:
- Kaplan-Meier analysis – estimate survival over time
- Log-rank test – compare survival curves between groups
- Cox proportional hazards model – assess effect of predictors on survival while adjusting for confounders.
- Competing risks methods – used when more than one type of event can occur
- Parametric survival models – model survival with assumed distribution
To check for the level of statistical significance, use the p-value from your test. You can report the actual p-values or use asterisks to denote…
To check for the level of statistical significance, use the p-value from your test. You can report the actual p-values or use asterisks to denote statistical significance. For asterisks, the following are commonly used: p < 0.05 = *; p < 0.01 = **; p < 0.001 = ***; p ≥ 0.05 = not significant.
For example:
Treatment A = 45*
Treatment B = 60
You can also use letters such as a, b, c when comparing statistical significance in multiple groups. Values with the same letter are described as not significantly different, while values with different letters are significantly different.
For example:
Group 1 = 10ᵃ
Group 2 = 12ᵃ
Group 3 = 18ᵇ
Running adjusted odds ratios (aORs) or adjusted risk ratios (aRRs) is a statistical approach used to control for confounding variables when assessing the relationship between…
Running adjusted odds ratios (aORs) or adjusted risk ratios (aRRs) is a statistical approach used to control for confounding variables when assessing the relationship between an exposure and a binary outcome. For example, smoking (yes/no) as the exposure and lung cancer (yes/no) as the outcome.
In Stata, adjusted odds ratios can be estimated using logistic regression, for example: logistic lungcancer smoking age sex bmi.
Adjusted risk ratios can be estimated using modified Poisson regression with robust standard errors, for example: glm lungcancer smoking age sex bmi, family(poisson) link(log) vce(robust) eform.
It is not ideal to use the Student’s t-test for a skewed distribution. This is because one of its key assumptions is that the data,…
It is not ideal to use the Student’s t-test for a skewed distribution. This is because one of its key assumptions is that the data, or more specifically the residuals, are approximately normally distributed. A skewed distribution can violate this assumption.
When this assumption is not met, alternative tests are more appropriate. These include the Mann-Whitney U test for two independent groups and the Wilcoxon signed-rank test for paired data.
