Urban canopy coverage has been linked with improved health outcomes, but most existing research is focused on canopy quantity rather than species composition. This study investigated whether canopy coverage predicts a composite health score across census tracts in Chicago, and aimed to motivate further research into the importance of urban tree species composition. IIT’s Alphawood Arboretum in Chicago was used as an empirical case study to lay the groundwork for future species specific analysis. CDC PLACES, ACS, and CRTI data were used to analyze health outcomes across 801 census tracts in Chicago. A composite health burden index was constructed using PCA, along with four progressive OLS regression models to predict the health burden index in different census tracts. Results indicated canopy coverage alone does not predict health burden (Model 1, p = 0.537), and the full model achieves \(R^2\) = 0.887, which is driven by the poverty rate, educational attainment, and racial composition. The IIT Alphawood Arboretum allowed for an analysis of 1442 trees across 77 different species, with a Shannon Index of \(H’ = 2.98\). The Thornless Honeylocust was the dominant species on campus, covering 32.0% of the campus, while also sequestering and storing the greatest amount of carbon. Priority planting tracts were also identified and were concentrated on Chicago’s South and West sides with Armour Square, South Chicago, and Fuller Park being the highest priority community areas. Neighborhood scale species composition data does not publicly exist for the city of Chicago and represents a genuine data infrastructure gap. The priority framework provides community organizations with evidence-based guidance on where to focus future investment. Future research should prioritize species-level documentation across Chicago’s urban forest to enable further testing of the species-health hypothesis.
Trees in urban areas have been shown to have various positive effects on public health. The connection has mostly been observed through canopy coverage. A denser urban forest increases canopy coverage, increasing overall shade and decreasing surface temperature. By observing differences in canopy coverage across neighborhoods of different demographics, McDonald, Biswas, Chakraborty, and others estimated that trees annually help avoid 632 ± 100 deaths in white neighborhoods and 442 ± 97 deaths annually in majority POC neighborhoods. This disparity can be attributed to a significant difference in tree canopy coverage between white neighborhoods and the majority POC neighborhoods [2]. This is significant to our research as there is a disparity in tree equity between neighborhoods in Chicago, which will be expanded on in the next section. Beyond heat-related complications, a denser urban forest can also be linked to a decrease in harmful fine particles such as (PM2·5), nitrogen dioxide (NO2), and tropospheric ozone (O3). A recent quantitative health impact assessment performed by Sicard, Pascu, Petrea, and others across 744 urban centres in 36 European countries estimated that for each five percentage point increase in tree canopy, 4,727 air pollution-related premature deaths could potentially be prevented annually across those cities [1]. Public health is not only concerned with physical health and mortality, however, as mental health is also a serious concern. The effect of canopy coverage on various mental health metrics has also been extensively studied. Canopy has been shown to have a positive impact in decreasing suicide attempts [19] and decreasing dementia rates [3]. Overall, prior research has indicated that higher canopy coverage can have positive effects on various public health metrics.
In 1934, the federal government set up the National Housing Act (NHA). This program was intended to encourage home ownership by offering low-cost loans. The risk that was associated with each of these loans was assessed, and neighborhoods in Chicago were split into four categories: A, B, C, and D. The Home Owners’ Loan Corporation (HOLC) was established the prior year to create security maps that labeled neighborhoods with these specific rankings. Race was explicitly used for the rankings, with A being deemed the safest, primarily white and US-born, newest structures, and D being deemed hazardous, mostly black and foreign-born people, with the oldest houses [5]. These assessments reflected racial assumptions rather than actual financial risk. On maps of the Chicago area, D neighborhoods were marked red, and this is where the term ‘redlining’ comes from.
Historical redlining has contributed to racial segregation and wealth disparities, which have shaped where public and private investment has been directed [6]. Redlined neighborhoods received reduced investment in green infrastructure over subsequent decades, and because urban canopy accumulates over time, these effects are still visible in the present landscape [7]. This pattern is evident nationally, as a 2021 study has found evidence of an association between worse historical HOLC grades and less green space in U.S. urban areas [6]. In Chicago, these neighborhoods still see the impacts of historic redlining, as A-ranked neighborhoods have about 27% more canopy coverage than D-ranked ones [5].
The Chicago Region Trees Initiative (CRTI) has extensively documented where canopy coverage is low, and the related challenges are amplified [8]. CRTI has developed maps of priority areas and conducted canopy assessments to improve equitable access to green spaces across Chicago. These maps are separated by neighborhood and ranked by the highest need for improved tree canopy, along with other variables such as surface temperature, air quality, and asthma risk [9]. There is a geographic concentration of low canopy coverage on Chicago’s South and West Sides, and our study aims to investigate how understanding species composition can help us better understand Chicago’s urban forest.
Research on Chicago’s urban forest has identified that species composition is a key determinant of ecosystem service delivery in urban forests. Large trees have greater per-tree benefits when it comes to pollution removal, carbon storage, and carbon sequestration. Additionally, using long-lived and low-maintenance trees offers the opportunity to reduce pollutant emissions from planting, removal, and maintenance activities. Evergreen trees should be used for their particulate matter benefits and because they offer year-round benefits compared to deciduous trees that lose their leaves when dormant [10]. These findings suggest that species selection is an essential consideration and deserves more attention as a dimension of the urban forest, because species identity shapes ecological value and human health outcomes.
Native tree species support significantly greater biodiversity when compared to invasive and non-native species. A study that compared lepidopteran (butterflies and moths) use of native plants versus introduced species found that native plants supported approximately four times more lepidopteran species than introduced plants. When the comparison was restricted to ornamental plants, native plants supported fourteen to fifteen times more lepidopteran species than introduced ornamentals [11]. Native tree species anchor broader food webs by supporting insect communities that other wildlife depend on. Insect diversity is an important consideration in urban areas, where non-native ornamental plantings outnumber native species. The native versus non-native species composition of Chicago’s urban forest is poorly documented and is a gap that this study aims to begin addressing through IIT’s Alphawood Arboretum.
Beyond ecological value and ecosystem services, species composition may affect physical and mental health. Bratman et al. (2019) identify species composition and biodiversity as potential factors affecting mental health [12]. Clearer links exist between urban tree cover and physical health outcomes. A 2008 study by Lovasi et al. concluded that street trees were associated with a lower prevalence of early childhood asthma; however, the causal relationship at the individual level remains to be established [13]. More direct evidence comes from a study documenting the spread of the emerald ash borer, which found that counties that experienced a loss of trees due to the emerald ash borer had increased mortality rates related to both cardiovascular and lower respiratory illness [14].
Species composition has been understudied compared to canopy coverage quantity. This study aims to contribute evidence supporting greater consideration of species composition, using IIT’s Alphawood Arboretum as a proof-of-concept for what species-level analysis can reveal at the neighborhood level in Chicago.
The long-term goal of this project is to come to data-driven conclusions about the effects of specific species, tree attributes, and biodiversity levels on different physical and mental health metrics in Chicago. In order to achieve this, a large sample of tree plots across a variety of neighborhoods or census tracts is necessary. Currently, these plots do not exist outside of college campuses and smaller, wealthier neighborhoods. To work around this, this report focuses on finding ways to best describe the data that does exist. In terms of large-scale Chicago data, this is in the form of census tract health and socioeconomic data and overall canopy coverage. In terms of species-specific tree plot data, this report will focus on IIT’s Alphawood Arboretum.
Prior research has contributed to finding the connection between canopy coverage and health benefits, and this study aims to extend these findings by incorporating species-level analysis. Our research will be driven by the following three questions:
Tree canopy coverage data were extracted from the CRTI. This dataset provides census-tract-level estimates of canopy coverage, vulnerability indices, surface temperatures, and air-toxin exposure for 801 census tracts. The CRTI priority scores are computed based on air toxin exposure, surface temperature, flood susceptibility, and social vulnerability. These data served as our primary source for urban tree canopy coverage in this analysis.
Chicago health outcome data were extracted from the CDC PLACES Local Data for Better Health, 2025 release. This dataset provides estimates for asthma, COPD, physical inactivity, mental health days, obesity, and depression at the census tract level. These data served as our primary source of neighborhood health outcome measures in this analysis.
Census tract boundary data were obtained from the US Census Bureau’s TIGER/Line Shapefiles, 2010 release. This dataset provides the boundaries for all 801 Chicago census tracts and was used to spatially join health, canopy, and socioeconomic data. Tract boundaries were matched based on a GEOID identifier.
Socioeconomic data were extracted from American Community Survey (ACS) 5-year estimate datasets, 2024 release. These datasets provided estimates about median income, poverty rate, education, and race/ethnicity at the census tract level in Chicago. These data were used as covariates to control for socioeconomic factors in the regression models.
Air quality data were obtained from the 2020 Chicago Air Quality and Health Index. This dataset provides air quality scores and percentile ranks for each census block group in the city. AQH scores were aggregated to the census tract level using population-weighted averaging prior to analysis.
Species-level tree inventory data were available for the Illinois Institute of Technology Alphawood Arboretum, located in the Douglas area (census tract 3510) on Chicago’s South Side. This inventory documents 1,442 established trees across 77 unique species, on IIT’s Mies campus. This dataset contained species name, common name, and diameter at breast height (DBH) measurement. These data were used for all analyses related to our case study of IIT’s Alphawood Arboretum.
A 2010 USDA Forest Service report [10] provided a table with pollution removal, carbon sequestration, and carbon storage estimates for trees based on their DBH class. These data were used to quantify the ecosystem service contributions of each tree species recorded in the IIT Alphawood Arboretum inventory.
All datasets were merged at the census tract level using the GEOID identifier. The ACS GEOID values had prefixes, so those numbers had to be truncated so that they matched the format of the TIGER/Line dataset. AQH scores existed at the block group level, and multiple block groups exist in each census tract, so block group scores were weighted by population before averaging to the tract level. The final merged dataset contained 801 census tracts, 31 columns, covering all of Chicago.
To clean the IIT tree inventory, 201 trees with implausible or missing DBH values were excluded, which can be attributed to data entry errors. Only established trees were retained, excluding any saplings that were recently planted.
Of 801 tracts, 39 (4.9%) were excluded from the regression for missing values, most commonly in one or more health or socioeconomic variables, AQH score, or median income. These tracts were concentrated on Chicago’s South and West sides, the same areas exhibiting the highest health burden and lowest canopy coverage. The analytic sample therefore under-represents the highest-need neighborhoods. This may bias the socioeconomic coefficients, though the direction was not explored in this study. These tracts were retained in the descriptive mapping but not the regression or ranking.
Five health outcomes were combined into a composite Health Burden Index to capture shared variance among the outcomes and reduce dimensionality. Asthma, COPD, physical inactivity, mental health days, and obesity were the five outcomes used, with higher values indicating worse health outcomes. Depression was excluded from the index because it was positively correlated with median income (r = 0.348, p < 0.001) and education (r = 0.460, p < 0.001), suggesting that this reflects access to healthcare and diagnosis rates rather than underlying health burden.
All five variables were expressed as census tract-level prevalence percentages. Before PCA, variables were standardized using z-scores. PC1 is the Health Burden Index, which summarizes a tract’s health across all five health outcomes simultaneously. Higher values on the Health Burden Index indicate greater health burdens (poorer overall census tract health). PC1 captures 86.1% of the variance, meaning that this one number explains almost all the information contained in the five separate measures. All five variables load positively, meaning that PC1 represents a unified health burden.
To validate this methodology, the PCA index was compared with the z-score average of all five health variables. A high correlation of r = 1.000 between the two confirms that the index is robust.
All regression models use census tracts as the unit of analysis. The composite Health Burden Index (PC1) is the primary response variable. Four progressive models were built rather than just running one model. This approach allows observation of how the canopy coefficient changes as additional factors are introduced. The models are as follows:
Bivariate — canopy only
+Environmental controls (AQH score, air toxins, surface temperature)
+Social vulnerability index
Full model + ACS socioeconomic controls (median income, poverty rate, educational attainment, percent Black residents, percent Hispanic residents)
Ordinary Least Squares regression was used because the response variable is continuous and the data were approximately normally distributed. The Breusch-Pagan test was used to test for heteroscedasticity, which means the variance of the residuals is not constant across predictors. Heteroscedasticity was identified in all four models, and to correct this, heteroscedasticity-robust standard errors (HC3) were applied. HC3 standard errors account for unequal residual variance and allow for more reliable inference.
For all predictors in Model 4, variance inflation factors (VIF) were computed, and most variables had a VIF below 5, except for the percentage of black residents (VIF = 8.73) and the percentage of Hispanic residents (VIF = 7.98). The results of this indicate that moderate collinearity exists between these two variables.
Consistent with the ecological fallacy, all of the results of this regression analysis represent tract-level ecological associations, and individual-level inferences cannot be drawn from this analysis. Moran’s I was computed to assess for spatial autocorrelation. A significant positive value (I = 0.291, p < 0.001) indicates that residual spatial autocorrelation remains after controlling for socioeconomic factors. A spatial regression model is recommended in any future work, and results should be interpreted under this known limitation.
To establish priority tracts, Health Burden Index scores were combined with canopy coverage data across all of the 762 census tracts with complete data. Tracts were ranked using a composite priority score, which was weighted equally between health burden rank and low canopy rank. This framework is intended to provide community organizations with evidence-based guidance on where targeted tree planting would have the greatest potential impact.
799 residential census tracts were analyzed to provide a better understanding of Chicago’s urban forest. The urban forest is characterized by a mean canopy coverage of 19.29%, a median of 18.79%, and a standard deviation of 7.84%. The lowest canopy tract is a census tract in the Edgewater area with a canopy coverage of 1.39%. Although Edgewater is a northside area, this specific census tract represents a dense high-rise residential area. The highest canopy tract is the Ohare residential area tract with a coverage of 69.04%.
CRTI’s priority tracts were also analyzed, and results showed that 312 tracts were classified as very high priority and 301 tracts were classified as high priority. CRTI’s analysis covered 782 of 799 residential tracts, because 17 tracts are missing rank values, but 78.4% of these were classified as high or very high priority. Very high priority tracts have a mean of 14.24% canopy coverage, whereas very low priority tracts have a mean canopy coverage of 39.54%. 73 tracts (9.3%) have canopy coverage below 10%. Areas of lower canopy coverage are concentrated in the downtown Chicago area, in and around the Loop neighborhood area. This represents the urban density and commercial land use in downtown Chicago.
Figure 1. Tree
Canopy Coverage Across Chicago Census Tracts
The five health outcome variables and citywide averages were asthma (10.52%), COPD (5.94%), physical inactivity (25.94%), mental health days (16.49%), and obesity (33.96%). The range of physical inactivity ranges from a minimum of 8.50% to 51.70%, which is the widest range of any of the health outcome variables. The wide range represents the substantial variation that the Health Burden Index is designed to capture.
Figure 2:
Health Burden Index across 782 residential census tracts in
Chicago
Areas of higher health burden are concentrated in Chicago’s south and west sides. This is consistent with the historical disinvestment that is discussed in section 1.2. The low canopy tracts in Chicago’s downtown do not correspond to high health burden. This suggests environmental and health inequity in Chicago’s South and West side, where there are elevated health burden scores.
PCA of the standardized health outcome variables yielded a first principal component (health burden index) that explains 86.1% of the variance. This confirms that a single component accurately represents the shared health burden across Chicago’s census tracts. All five health outcome variables load positively, ranging from 0.432 (asthma) to 0.470 (obesity), as shown in Figure 3. This represents a unified health burden dimension.
Figure 3: PCA
Construction of the Health Burden Index
All four models are presented in Table 1, showing \(R^2\), Adjusted \(R^2\), canopy coefficient, and significance.
Table 1:
Regression results across four models
Model 1 is the bivariate result, and canopy alone is not significant at the α = 0.05 level (p = 0.5375). Model 2 adds environmental controls (AQH score, air toxins, and surface temperature), and model 3 also adds the CRTI social vulnerability index. For these models, the canopy coefficient is positively and significantly associated with the composite health burden index, where model 2’s p-value is 0.0114 and model 3’s is 0.0054. This is not a causal relationship and represents confounding between the variables. Wealthier neighborhoods tend to have higher canopy coverage and better air quality, so once air quality is controlled, canopy coverage appears to have a positive association with health burden. Model 4 is the full model, and it adds poverty rate, educational attainment, and racial composition. In this model, the canopy coefficient is no longer significant (p = 0.1747).
The full model has an \(R^2\) of 0.887, with the variance predominantly explained by poverty rate, educational attainment, and racial composition, as shown in Figure 4. The significant predictors are poverty rate, percentage of residents with bachelor’s degrees, air quality score, and air toxins, as illustrated in Figure 5. HC3 robust standard errors were applied, but this did not change the conclusions. Moran’s I score was computed on model 4’s residuals (I = 0.291, p < 0.001), which confirms significant spatial autocorrelation. Neighboring census tracts cannot be treated independently, so results should be interpreted with this limitation in mind.
Figure 4: Progression of \(R^2\) through all four models
Figure 5: Model
4 coefficient plot with 95% confidence intervals from robust standard
errors
The priority tracts were identified using a formula that assigns 50% weight to the health burden rank and 50% to the canopy coverage rank across the 762 census tracts with complete data. The highest priority tracts based on this analysis are in the community areas, Armour Square, South Chicago, and Fuller Park. The top 10 priority census tracts have an average canopy coverage of 8.38% with health burden scores ranging from 2.949 to 6.06. Health burden scores are PCA-derived and are expressed as standard deviations away from the mean = 0. Priority planting is geographically concentrated in Chicago’s west and south sides, as illustrated in Figure 6.
Table 2: Top 10
priority planting census tracts
Figure 6: The
priority planting map of Chicago
Canopy coverage alone does not significantly predict neighborhood health outcomes in Chicago. When socioeconomic factors are controlled, the relationship between canopy coverage and composite health burden disappears. \(R^2\) rose from near 0 in the bivariate model to 0.601 when environmental controls were introduced and jumped again to 0.887 when socioeconomic controls were introduced to the model. This shows that socioeconomic controls such as poverty rate, educational attainment, and race are dominant factors in health burden.
A notable pattern becomes clear in Models 2 and 3, where controlling for air quality causes canopy coverage to become a significant predictor. This is because wealthier, healthier neighborhoods tend to have both more canopy coverage and better air quality. The remaining variation of the model in the canopy is driven by neighborhood wealth rather than individual tree species. The ecological fallacy applies here because all results are tract-level associations and cannot be used to draw any conclusions on individuals in these census tracts. Tracts with more canopy coverage do not necessarily indicate that the residents have better health. Results should be interpreted with the spatial correlation limitation in mind, as mentioned in section 3.3.
These results reinforce the importance of species composition as a research focus, instead of just tree coverage quantity. The null canopy result does not mean that trees are irrelevant to health, but instead shows that census-level canopy coverage data is too confounded and coarse to draw species-specific conclusions. Past research has linked street trees to lower asthma prevalence [13] and a loss of trees linked with increased cardiovascular and respiratory mortality [14]. With more precise data, further analyses can be conducted so that specific trees can be identified that have the greatest positive impact on health.
The priority planting framework suggests investment into priority areas that would see the biggest improvements in health burden and canopy coverage. The top ten priority tracts have been identified in Table 2, and some priority areas include Armour Square, South Chicago, Fuller Park, Greater Grand Crossing, and Washington Park. Based on the Alphawood Arboretum inventory, the Thornless Honeylocust is the dominant species, which also sequesters and stores the most carbon. Trees like this that have performed well in a South Side environment should be considered for further planting across priority census tracts, especially across the South Side.
Urban forestry best practices for protecting against pests suggest that an urban forest should contain no more than 10% of any single tree species, 20% of any single genus, and 30% of any single family [17]. The Thornless Honeylocust makes up 32% of the trees on IIT’s campus which is more than three times the suggested 10% species guideline as recommended by the 10/20/30 rule. Further planting in the Alphawood Arboretum should keep this figure in mind and should aim to increase biodiversity.
For targeted planting across Chicago, several actionable recommendations can be made. Prioritizing species that reach large DBH, have long lifespans, and require minimal maintenance will deliver the greatest long-term ecosystem benefits [10]. Additionally, high-benefit trees that have been successfully planted and maintained in certain geographic areas should be further utilized in areas with low canopy coverage and high health burden. In order to quantify benefits specific species bring, existing tree inventories and any future trees planted should be documented for further analysis
To perform a standalone analysis on IIT’s Alphawood Arboretum, we needed to determine what was possible to analyze with our data, what aspects of the urban forestry we wanted to focus on, and which mathematical analysis methods would be necessary to achieve these objectives. Our objectives for the analysis were:
In order to quantify biodiversity, we used Shannon’s Diversity Index \((H' = -\sum_i p_i \ln p_i)\), where \(S\) represents species richness, and \(pᵢ\) represents the proportional abundance of each species \(i\). Pielou’s evenness \((J' = \frac{H'}{\ln(S)})\) was also calculated to decouple contributions of richness and evenness to overall diversity.
The arboretum’s ecosystem services were quantified using the i-Tree Eco v6.040 model [16], applied to field inventory data collected at Alphawood in 2026 [18].
i-Tree, an urban and rural forestry analysis assessment tool provided by the USDA Forest Service, was used to assess forest structure and quantify environmental effects due to Alphawood’s urban forest. To run the model, the arboretum dataset was input with columns for tree ID, DBH in inches, tree species scientific name, and latitude/longitude. Before running the analysis, the location for Chicago, Cook County, Illinois, was entered. The Midway Airport pollution and weather data collection centers were selected for the model to estimate location-specific pollution and air quality levels.
i-Tree utilizes species-specific equations to calculate dry biomass from provided DBH values. If no species-specific equation exists, the model falls back to a genus-level equation, then to a hardwood or conifer average. Open-grown urban trees tend to be less dense than forest trees, so their biomass was reduced by 20%. Biomass was then converted to carbon by multiplying by 0.5. In order to estimate sequestration, the i-Tree model added average annual diameter growth to the current diameter to estimate what the tree would store one year later, estimating annual carbon sequestration. The i-Tree model has the ability to calculate leaf area via crown dimensions and percentage of crown canopy missing. Since the Arboretum data did not contain this information, the model estimated it based on species and DBH values. This was used to estimate tree cover and leaf area values for each tree plot. Since tree crown health was also not measured, the model used a default 13% dieback, or 87% condition. This means that the growth rate and health of dead or dying trees would be overestimated, which must be taken into consideration in the analysis.
Using i-Tree’s hybrid big-leaf and multi-layer canopy deposition model, hourly tree-canopy resistance values were calculated for ozone, sulfur dioxide, and nitrogen dioxide. Carbon monoxide and particulate matter removal used average measured values from literature that were adjusted based on leaf phenology and leaf area, and a 50% resuspension rate of particles back into the atmosphere was incorporated in the particulate removal estimates.
i-Tree Eco also produces monetary valuations for pollution removal and carbon (pollution valued via avoided-damage/externality costs, carbon via a social-cost-of-carbon rate). These are model-generated estimates of societal avoided-damage value, not healthcare or public-spending savings, and are reported for context and not analyzed quantitatively.
It is important to address that these results are estimates dependent not only on the quality of the data collected, but also on the quality of weather and pollution data collection centers in the area. The nearest data collection center for both weather and pollution data was the Midway Airport location, which was labeled as “poor” quality by the i-Tree Eco software at the time of processing. Furthermore, the newest version of the software recommends inputting the crown measurements described earlier, which were not measured in the arboretum. The effects of these parameters will be discussed.
The Alphawood Arboretum contains 1643 tree plots across the campus, comprising 82 unique species. However, 201 of the 1643 plots had missing or implausible DBH values. After dropping those entries for the analysis, we are left with 1442 trees over 77 unique species. Of these 1442 plots, Thornless Honeylocusts made up 32.0% of the total, followed by Hawthorn, making up 11.0%, and Green Ash, 6.7% of the total.