Biostatistics Guide: Relative Risk Formula & Study Designs
Biostatistics Guide

A Comprehensive Biostatistics Guide for PG Residents 

Biostatistics Guide

Welcome to your essential biostatistics review designed specifically for PG residents. Mastering research methodology and statistical analysis is a core competency during your residency. From interpreting standard deviations to calculating outcomes using the relative risk formula, this guide breaks down high-yield concepts into digestible, factually accurate sections. 

Get High Quality, High Yield PG Residency Notes for Free! 

Understanding Observational Study Designs 

Classifying research based on the temporal relationship between exposure and outcome is fundamental.  

  • Cross-Sectional Study (Present): This design acts as a snapshot or survey of a population. It involves the simultaneous measurement of exposure and outcome to determine the prevalence percentage of a condition. An example includes measuring mean blood pressure in three communities via house visits.  
  • Case-Control Study (Past): This study starts with the outcome by looking at cases (diseased) and controls (non-diseased). It then traces back to find the exposure history. For instance, comparing DVT patients versus controls to find past risk factors.  
  • Cohort Study (Future): This study starts with exposure (e.g., smokers vs. non-smokers) and follows the groups forward to observe disease development. This design relies heavily on the relative risk formula to quantify outcomes. An example is following pregnant smokers to see if they deliver low birth weight babies.  

The 2×2 Table and the Relative Risk Formula 

Quantifying the link between exposure and disease requires utilizing the standard 2×2 table. This matrix is the foundation for applying the relative risk formula and calculating odds ratios.  

The 2×2 Matrix for Measuring Risk 

Status Diseased (Cases) Non-Diseased (Controls) 
Exposed a  b  
Non-Exposed c  d  

In case-control studies, we use the Odds Ratio (OR), calculated as OR = (a × d) / (b × c). For example, if OR = 6, it implies exposed individuals have a 6x higher chance of having the disease. 

However, in cohort studies, researchers apply the relative risk formula. The relative risk formula measures the incidence in the exposed group divided by the incidence in the non-exposed group. 

Mathematically, the relative risk formula is expressed as: RR = (a / (a + b)) / (c / (c + d)) 

By using the relative risk formula, you can determine the exact risk multiplier.  

Interpreting the Results of the Relative Risk Formula 

Once you calculate the data using the relative risk formula or the OR, interpretation follows strict rules:  

  • Value = 1: No association (Null).  
  • Value > 1: The exposure is a risk factor and increases disease likelihood.  
  • Value < 1: The exposure is a protective factor and prevents disease.  

What is the difference between “Intention to Treat” and “Per Protocol” in Randomized Control Trials? 

“Intention to Treat” includes all drop-outs in the final analysis, making it the preferred method as it reflects real-world conditions. Conversely, “Per Protocol” excludes drop-outs and evaluates the drug only under ideal conditions.  

Efficacy and Number Needed to Treat (NNT) 

In Randomized Control Trials (RCTs), researchers analyze a treatment group against a control group (placebo). The efficacy of a treatment can also be linked back to our previous risk measurements. Efficacy is calculated as 1 – Relative Risk (RR). You must use the relative risk formula accurately first to ensure your efficacy calculation is correct. 

Another vital metric is the Number Needed to Treat (NNT). NNT represents the minimum number of patients that must be treated to achieve one additional cure compared to the control group. 

The formula is: 
NNT = 1 / Absolute Risk Reduction 
NNT = 1 / (Incidence in Control – Incidence in Treated) 

If the control cure rate is 10% (0.1) and the treatment cure rate is 20% (0.2), the gap is 0.1. Therefore, NNT = 1 / 0.1 = 10, meaning you must treat 10 patients to benefit 1 extra person. 

Evidence Synthesis and Reporting Standards 

When combining data through a Meta-Analysis, systematic combinations of results from previous studies are used to reach new conclusions (data is collected from studies, not directly from patients). Reporting these findings requires strict adherence to guidelines.  

Research Study Reporting Guidelines 

Study Design / Focus Guideline Acronym 
Systematic Reviews & Meta-Analyses PRISMA (Preferred Reporting Items)  
Meta-Analyses (Quality Focus) QUOROM (Quality of Reporting)  
Randomized Control Trials (RCT) CONSORT (Consolidated Standards)  
Observational Studies STROBE (Strengthening Reporting)  
Diagnostic Accuracy Studies STARD (Standards for Reporting)  

Diagnostic Testing and Predictive Values 

Evaluating diagnostic tests is another pillar of biostatistics. 

• Sensitivity: Defined as TP / (TP + FN), it is the true positive rate used for screening. The goal is to find all diseased cases. 

• Specificity: Defined as TN / (TN + FP), it is the true negative rate used for confirmation. The goal is to rule out the non-diseased. 

After a test result is known, we look at post-test probability using Predictive Values. 

• Positive Predictive Value (PPV): TP / (TP + FP). This is the probability that a positive patient actually has the disease. 

• Negative Predictive Value (NPV): TN / (TN + FN). This is the probability that a negative patient is actually disease-free. 

Hypothesis Testing and Data Distribution 

When conducting research, understanding errors is critical. 

• Type I Error (Alpha): A false positive where you reject the null hypothesis when it is actually true. The P-Value measures the probability of committing this error. 

• Type II Error (Beta): A false negative where you accept the null hypothesis when it is false. Power is the probability of correctly finding a difference, calculated as 1 – Beta. 

Measures of Central Tendency and Skew 

Data is analyzed using the mean (sensitive to extremes), median (robust to outliers), and mode (most frequent value). In a perfectly Normal Distribution (Bell Curve), the Mean, Median, and Mode are all equal. 

The empirical rule dictates that 68% of data falls within Mean ± 1 SD, 95% within Mean ± 2 SD, and 99.7% within Mean ± 3 SD. 

However, when data is asymmetrical, it creates a skew:  

  • Right Sided (Positive) Skew: The mean is dragged by a high outlier, resulting in Mean > Median > Mode.  
  • Left Sided (Negative) Skew: The tail extends to the left, resulting in Mean < Median < Mode.  

Frequently Asked Questions (FAQs) 

1. What is the primary use of the relative risk formula?  

The relative risk formula is primarily used in cohort studies to quantify the incidence of disease in an exposed group versus a non-exposed group.  

2. How do you calculate Odds Ratio (OR)?  

Using a 2×2 table, the OR is calculated as (a × d) / (b × c), mostly utilized in case-control studies. 

3. Does the relative risk formula determine prevalence?  

No, a cross-sectional study determines prevalence through a simultaneous snapshot survey. The relative risk formula measures risk moving forward in time.  

4. What does an RR value of less than 1 indicate?  

When applying the relative risk formula, a value less than 1 indicates a protective factor, meaning the exposure helps prevent the disease.  

5. What is the formula for the Number Needed to Treat (NNT)?  

NNT is 1 divided by the absolute risk reduction, or 1 / (Incidence in Control – Incidence in Treated). 

6. Which reporting guideline is used for Randomized Control Trials?  

The CONSORT (Consolidated Standards) guideline is utilized for reporting Randomized Control Trials.  

7. When should I look at Sensitivity vs. Positive Predictive Value (PPV)?  

Use Sensitivity before a test is run to understand its screening characteristics, and use PPV after the result is known to determine patient probability.  

8. What happens during a Type I Statistical Error?  

A Type I error is a false positive where researchers incorrectly reject the null hypothesis when it is actually true.  

9. How does an outlier affect the mean and median?  

An extreme outlier will drastically drag the mean toward it, while the median remains robust and unaffected, making it a better measure for skewed data.  

10. What defines a Right Sided (Positive) Skew?  

A right-sided skew occurs when the tail extends to the right, creating a sequence where the Mean is greater than the Median, which is greater than the Mode. 

Latest Blogs

PG Residency 2+1 Plan