Beyond State Lines: The Hidden Forces in County Housing Prices
Principal components let historical housing behavior—not a rigid political boundary—describe the state influence on county markets.
Our previous analysis found that county housing results clustered by state. State dummy variables improved the regression, but they did so with a blunt instruction: a county is either inside a state or outside it. Idaho equals one in the Idaho column and zero everywhere else; New York receives the reverse treatment. That identifies a boundary, but it says nothing about how the two housing markets actually behave.
We wanted the data to describe those relationships. We therefore replaced the state indicators with principal components calculated from the quarterly growth of house-price indexes for all 50 states and the District of Columbia from 1975 through 2023.
Three patterns emerge
Principal component analysis compresses many correlated state histories into a smaller set of uncorrelated national and regional patterns. The first component captures the largest shared movement. The second captures the largest remaining pattern after removing the first, and so on.
PC₀ — The common national housing cycle
Selected state loadings. Every state loading is positive; the difference is intensity.
The first component is unmistakable: it is the broad national housing cycle. Interest rates, national credit conditions, construction costs, migration and common expectations can move markets together even when their local economies differ.
PC₁ — Differences in state-cycle timing
Selected extremes; blue is positive and red is negative. A component's overall sign is arbitrary.
The second component is not a clean map of East versus West or urban versus rural. It appears to capture differences in timing, volatility and sensitivity across state housing cycles. That is still valuable: PCA is allowed to discover a relationship that does not conform to a familiar political or geographic label.
PC₂ — Northeast versus energy and intermountain markets
Selected extremes reveal the clearest regional contrast among the first three components.
The third component is easier to name. Northeastern states load strongly on one side, while Texas, Louisiana, Oklahoma, Wyoming, Colorado, Idaho and Utah tend toward the other. The component appears to distinguish the Northeast from energy-producing and intermountain housing markets.
From state history to county regression
Our county cross-section contains 213 complete observations from Washington, Oregon, Idaho, Texas, New York, Illinois, Maryland and West Virginia. The outcome is county house-price growth. The local explanatory variables are personal-income growth, listing growth and nominal GDP growth. We then give every county the first three PCA coordinates of its state.
This does not erase state boundaries completely: counties in the same state still receive the same state coordinates. It does, however, replace unrelated state labels with continuous measures of historical similarity. New York and Idaho no longer differ merely because their dummy columns say so; they differ according to how their housing histories load on the common factors.
| Variable | Coefficient | Standard error | t-statistic |
|---|---|---|---|
| Constant | 0.064 | 0.009 | 7.510 |
| Personal-income growth | −0.042 | 0.154 | −0.273 |
| Listing growth | −0.004 | 0.008 | −0.490 |
| Nominal GDP growth | 0.021 | 0.063 | 0.338 |
| PC₀: common national cycle | −0.110 | 0.047 | −2.310 |
| PC₁: state-cycle timing | 0.045 | 0.010 | 4.451 |
| PC₂: regional contrast | 0.080 | 0.009 | 8.607 |
What the t-statistics tell us
The three historical state factors are statistically much stronger than the contemporaneous county measures in this specification. PC₂ has a t-statistic of 8.61, PC₁ reaches 4.45 and PC₀ reaches −2.31. By contrast, personal income, listings and nominal GDP have t-statistics close to zero.
This does not prove that a principal component causes house prices. A component is a compressed description of correlated housing history, not a policy lever. The results do show that broad national and regional housing structure provides far more explanatory power for this 2024 county cross-section than the three local growth measures considered here.
We tried adding more components
For curiosity, we expanded the regression. With seven components—the maximum identifiable number for eight states after including an intercept—R² increased to 0.451 and adjusted R² to 0.424. But no later component produced a t-statistic above two: PC₃ reached 1.84, while PC₄ through PC₆ were weak.
A ten-component regression was singular. That failure is not a software mystery; it is matrix algebra. Eight states provide only eight distinct state profiles, and the intercept consumes one dimension. At seven state components, the PCA representation is already approaching the same state space as a complete set of state dummies. The three-component model is therefore the more meaningful result: it achieves substantial explanatory power while genuinely reducing dimension.
Housing markets do not stop at lines on a map. Their common national cycle and regional histories carry information that a collection of unrelated state labels cannot express.
What comes next
This remains an exploratory cross-section, not the last word on American housing. It uses eight states, one county observation date and three local covariates. Further work should include all states, multiple cross-sections, mortgage rates, construction constraints, migration and measures of local housing supply elasticity. The PCA method gives that work a disciplined starting point: preserve what state histories have in common, retain the regional contrasts, and avoid pretending every border creates an entirely separate market.