What it is (and is not)
The one claim
Section titled “The one claim”LookAlike is a discovery and triage tool for retail real estate. You pick a US retail brand; it profiles the neighborhoods where that brand’s stores already operate; then it ranks every continental-US census tract by how closely its observable characteristics resemble that archetype—and shows you, per neighborhood, exactly which features drove the match.
That is the honest claim. It is not a sales predictor. It is a way to decide where to look, not a recommendation to open.
What it does
Section titled “What it does”Pick a brand. A searchable picker at the top of the left panel lists eleven chains (Trader Joe’s, Whole Foods Market, Sprouts Farmers Market, Chipotle Mexican Grill, Starbucks, Costco Wholesale, Dollar General, Tractor Supply Co., Sweetgreen, H Mart, Cracker Barrel) with store counts sourced from OpenStreetMap — or add any other United States chain from OpenStreetMap by name or Wikidata identifier and score it live in the browser (with an OpenStreetMap-coverage caveat). Brands you add are saved in your browser and reappear in the picker under “Your brands” on later visits. Switching among brands you have already viewed is instant, because each brand’s national scores are cached in memory rather than re-computed.
See a national resemblance map. The map colors only high-similarity tracts (the median tract is left dark on purpose, so you see where to look, not just noise). Click any neighborhood to see:
- A similarity percentile (a within-brand rank from 0–100, where ~100 is the closest tract in the country to that brand’s archetype, and 50 is the median)
- An absolute-fit band (Strong, Moderate, or Weak) measured against the brand’s own store neighborhoods, which — unlike the percentile — is comparable across brands
- A low ACS reliability badge when the tract has a small population or a large income margin of error
- A radar fingerprint showing how that tract compares to the brand archetype across all 21 features
- A “why it matches” breakdown (the features it’s closest to the archetype on)
- A “where it differs” section (the features where it diverges, with the brand’s typical value for each)
- A feature-comparison table grouped by demographic theme
- A plain-English assessment of the trade-offs
Filter and refine. Open markets only (tracts without the brand already), by state, by metro area (choose from 383 Metropolitan Statistical Areas), by minimum similarity threshold, or by competition (at least a chosen distance from your own stores, or no tracked competitor within 6 km).
Rank metros. A “Top metros” view ranks metropolitan areas by how many open, high-resemblance-rank neighborhoods each holds — a starting point for “which market first?”, stated honestly as rank concentration, not market quality.
Overlay existing stores. See where the brand currently operates and the distance from any candidate to the nearest store (out to 6 km).
Upload your own stores. Build a custom archetype from your own store list. Authenticated CSVs become private, versioned retail-client datasets with persisted tract assignments and metric definitions; national scoring remains live in the browser. Optional columns let you weight toward better-performing locations or avoid lookalikes of failed, closed, or underperforming locations.
Tune the factors. Reweight the demographic themes (income/wealth/housing, education, density/transportation, age/household, race/ethnicity, retail/dining context) or switch themes on and off entirely. Each theme also opens an advanced per-feature layer to weight a single feature within a group. The map and rankings re-score live. These weights are your scenario assumptions, surfaced transparently.
Compare two brands. See a diverging map (a fixed, colorblind-safe blue-and-orange pair shows which brand a tract leans toward; brightness shows how strongly it resembles either), plus three ranked lists (Leans A / Resembles both / Leans B). A custom upload can also be compared against a built-in brand.
Check robustness. For any tract, ask whether its high rank is robust or hinges on one theme. The tool re-scores it under equal weighting and with each theme removed.
See nearby chains. Which of the tracked brands have a store within 6 km—a quick competitive-density read.
Mark “not a fit.” Flag a lookalike that is wrong on the ground; the tool hides it, and once you have flagged a few, it shows the features your rejected tracts most share—teaching you where the method misleads.
Shortlist and export. Star neighborhoods into a persistent shortlist (saved per brand in your browser), add notes, and export a print-ready PDF report or CSV (with the resemblance disclaimer, the fit band, and the reliability flag baked in). The ranked-list CSV export also carries the competition columns.
Save and share. Persist named scenarios (brand, weights, compare, metro, filters) to your browser to switch between markets, and use the Copy link button to put the current shareable view on your clipboard. The URL restores the brand, theme weights, filters, metro, compare brand, selected tract, and pinned shortlist; per-feature (Advanced) weights and custom store uploads live only in your browser, so they travel through named scenarios rather than the link.
How a match is computed (honest version)
Section titled “How a match is computed (honest version)”Every continental-US census tract (~83,000) is described by 21 features: demographics and housing from the US Census American Community Survey (median income, education, density, age, household composition, tenure, vehicle access, commute mode, race/ethnicity, poverty) plus retail context from Overture Maps (density of all points of interest, grocery stores, restaurants, cafés, and shopping per square kilometer).
A brand’s store locations are mapped to the tracts they sit in. The archetype is the average of those tracts’ features—one observation per occupied tract, so a Costco gas pump and warehouse don’t double-count.
Each feature is scaled by how tightly the brand clusters on it (a standardized deviation) with a per-feature variance floor (λ = 0.5): the floor raises the effective spread of a feature the brand is artificially narrow on, so it can’t overwhelm the rest, and never reduces the spread of a feature the brand is already wider on than the nation. Values are clipped to the national 1st–99th percentile first, so a few extreme tracts (top-coded incomes, ultra-dense cores) can’t distort the scale.
To prevent correlated features from being double-counted (income, per-capita income, home value, and rent are all r > 0.7), the 21 features are organized into 6 thematic groups. The distance averages within each group, then across groups, so each theme counts once. The affluence cluster can’t act as several separate votes.
Distance becomes a 0–100 similarity percentile—a within-brand rank among all scoreable tracts. ~100 is the closest tract in the country to that brand’s archetype; 50 is the national median. Percentiles are NOT comparable across brands as absolute fit strength.
Validation: Does the archetype actually mean anything?
Section titled “Validation: Does the archetype actually mean anything?”For each brand, the tool reports an archetype coherence score from k-fold footprint recovery: we rebuild the archetype with some stores held out, then check where those held-out store neighborhoods rank.
- Sweetgreen recovers its own held-out stores to about the 90th percentile—its locations really do share a recognizable profile. Resemblance means something.
- Starbucks recovers to about the 57th percentile—it operates in so many kinds of neighborhood that its archetype is diffuse and resemblance is a weak signal.
The folds are blocked by county, so a held-out store’s same-county neighbors cannot leak into training and inflate the number. This is an honest validation check, not a sales claim. It tells you how much weight to put on the results for each brand.
Data sources (all open, no API key required)
Section titled “Data sources (all open, no API key required)”| Layer | Source | License |
|---|---|---|
| Demographics | US Census ACS 2023 5-year, table-based Summary File (bulk .dat) | Public domain |
| Tract boundaries + land area | US Census TIGER/Line 2024 | Public domain |
| Retail context (POI density) | Overture Maps Places (public GeoParquet on S3, release 2026-06-17.0) | CDLA-Permissive 2.0 |
| Metro areas | US Census CBSA delineation (OMB 2023); 383 Metropolitan Statistical Areas | Public domain |
| Brand store locations | OpenStreetMap via the Overpass API (brand:wikidata / exact name) | ODbL |
| Basemap | CARTO dark (no-token raster) over OpenStreetMap | CARTO / ODbL |
See data sources for exact URLs and tables.
What it is NOT (scope boundary)
Section titled “What it is NOT (scope boundary)”Not a revenue, sales, or foot-traffic forecaster. A neighborhood’s resemblance to your brand’s archetype says nothing about how much money it will make.
Not a recommendation engine. It does not say “open here.” It says “this place looks like your winning markets; go investigate.”
Not a substitute for due diligence, brokerage, or local knowledge. Real estate availability, zoning, rent, competition dynamics, operations, and local relationships all matter more than resemblance.
Not a black box. Every match is feature-driven and explained. You see exactly which 3–5 factors drove the similarity.
Not a gentrification or displacement forecaster. The tool may surface neighborhoods undergoing demographic change. Using it to gentrify communities is an application choice, not the tool’s intent.
Not dependent on proprietary data. The tool ships entirely on open data to demonstrate that the weak-correlation problem in site selection is not a data-access problem—it is a fundamental fact about geography.
Honest limitations
Section titled “Honest limitations”- OpenStreetMap store coverage is good for large national chains but imperfect; a few stores may be missing or stale. Upload your own store list for an exact footprint.
- ACS carries a multi-year lag and sampling error, worst in small tracts. 2023 data reflects 2018–2023 conditions, with 5-year averaging.
- The feature set and weighting are interpretable choices, not ground truth. Different reasonable features would shift rankings somewhat. Trust the explanation over the exact rank.
- Resemblance ignores crucial real-world factors: rent, real-estate availability, zoning, competition dynamics, operations, and local knowledge—all of which actually decide whether a location works.
Who it’s for
Section titled “Who it’s for”Franchise developers vetting new markets before scouting in person.
Tenant-rep brokers doing early market screening for restaurant and retail clients.
Retail real-estate teams looking for a discovery tool to surface candidate neighborhoods and spark due-diligence conversations.
Anyone curious where a brand’s “kind of neighborhood” recurs across the country.
Next steps
Section titled “Next steps”- How the scoring works—the detailed method, the 21 features, and how to read a robustness badge
- Using custom data—how to upload your own stores and what happens when you weight by performance
- Tuning the weights—what scenario you’re asking the tool to answer
- Comparing brands—what a diverging map means and how to interpret the lean rankings