12 research questions for a future classroom or capstone library. They are not assignments currently in the field. Projects involving human participants are marked for likely IRB review.
Public policycapstone · advanced
Can regulatory enforcement predict subsequent consumer complaints?
Question. After a public enforcement or disciplinary event appears in a regulator’s records, do later consumer-complaint patterns change in a measurable way — and in which verticals is the relationship weakest?
Why it matters. Consumers and journalists often treat a single enforcement record as a lasting risk signal. Measuring lag, attenuation, and missing-data bias would test that assumption without implying TrustHub scores.
Verticals. Move Trust Hub, Lender Trust Hub, Contractor Trust Hub
Potential output. Working paper plus a replication notebook using a future versioned snapshot — not a TrustHub ranking product.
Public policymasters · substantial
Geographic variation in regulatory enforcement intensity
Question. How much of the observed difference in published enforcement counts across U.S. states is explained by industry size versus publication practice versus genuine enforcement intensity?
Why it matters. State boards and DOIs do not publish equally. Treating raw counts as “risk maps” can punish transparency rather than misconduct.
Verticals. Contractor Trust Hub, Insurance Trust Hub, SeniorTrustHub
Potential output. State-comparison memo with a documented incompleteness appendix.
Data science / AIcapstone · advanced
Measuring entity-resolution precision and recall across public business registries
Question. Using a future hand-reviewed TrustHub Entity Resolution Benchmark, what precision, recall, and false-positive rate would a simple name-and-place matcher achieve versus a more conservative exact-identifier matcher?
Why it matters. Public registries duplicate, rename, and DBA-shift. Over-merging unrelated firms is a consumer-protection failure, not a convenience.
Verticals. Cross-network
Potential output. Benchmark card and error taxonomy. No claimed TrustHub performance in Academic 001A.
Data science / AIcapstone · advanced
Phoenix-company detection in public records
Question. Can public regulatory records identify operators that dissolve or abandon an identity and later reappear under a new business name, address cluster, or officer set — without claiming a legal finding of fraud?
Why it matters. Household-goods moving and some construction markets have a documented public-interest problem of successive identities. Measurement should stay evidence-based and cautious.
Verticals. Move Trust Hub, Contractor Trust Hub
Potential output. Method paper emphasizing uncertainty, with withheld names unless counsel clears a public version.
Data science / AImasters · substantial
Calibration audit of TrustHub-derived research signals (where they exist)
Question. Where a specialist hub publishes an explicit research score or composite signal, how well do those values calibrate to later independently observed outcomes — and where a hub does not publish a proprietary Trust Score, what happens if researchers invent one anyway?
Why it matters. Not every hub uses a proprietary Trust Score. Forcing a universal score across moving, lending, insurance, contractors, senior care, and investment firms would misstate the methodology.
Verticals. Cross-network
Potential output. Hub-by-hub methods note: scored vs unscored evidence models.
Public policyundergraduate · moderate
Regulatory transparency across U.S. states: licensing and disciplinary records
Question. How accessible are licensing lookups and disciplinary records for contractors and insurance producers across states — by search UX, bulk data, and historical retention?
Why it matters. Consumer protection depends on whether the public can actually retrieve the record, not only whether a statute exists.
Verticals. Contractor Trust Hub, Insurance Trust Hub
Potential output. Open codebook and state scorecard of publication practice — not of “bad actors.”
Data science / AImasters · substantial
Missing-data bias in consumer regulatory datasets
Question. When a field, county, or state is systematically missing from public extracts, which research conclusions flip if missingness is treated as “clean” versus “unknown”?
Why it matters. TrustHub’s public stance is that missing evidence does not mean a clean record. That claim is testable.
Verticals. Cross-network
Potential output. Methods appendix suitable for later dataset dictionaries.
Data science / AIundergraduate · moderate
Business-name normalization and DBA matching
Question. How much of apparent “duplicate company” volume in public registries is explained by punctuation, legal-suffix, and DBA variation versus genuinely distinct entities?
Why it matters. Name cleaning is the cheapest matcher and the easiest way to create false merges.
Verticals. Move Trust Hub, Lender Trust Hub, InvestorTrustHub
Potential output. Normalization spec that a later academic snapshot can cite.
Public policycapstone · advanced
Ownership networks and nursing-facility regulatory outcomes
Question. Where CMS publishes ownership or chain relationships, how do inspection or enforcement patterns cluster — and what cannot be claimed because ownership files are incomplete or lagged?
Why it matters. Nursing-facility ownership is a live public-policy subject. Research must stay inside CMS (and supported state) evidence, not placement marketing.
Verticals. SeniorTrustHub
Potential output. Policy brief with a limitations box on CMS lag and unpublished state actions.
Data science / AIindependent study · substantial
Measuring regulatory-data freshness and historical information loss
Question. When public systems overwrite prior license status, how much historical information is lost between two calendar dates — and what research designs become impossible without immutable snapshots?
Why it matters. A historical snapshot can become more valuable with time precisely because regulators overwrite. This project motivates the versioning standard without running a production snapshot job.
Verticals. Cross-network
Potential output. Protocol for longitudinal academic releases (see dataset-release standard).
Data science / AImasters · advanced
Algorithmic confidence calibration for business identity matching
Question. If a matcher emits a confidence score, does predicted confidence equal empirical match frequency across bins — or are high-confidence errors concentrated in DBA and successor cases?
Why it matters. Uncalibrated “95% match” labels are how false merges get into consumer-facing products.
Verticals. Cross-network
Potential output. Calibration report template for later Academic 001B+ benchmark files.
Public policycapstone · substantial · IRB likely
Information presentation and consumer decision quality
Question. When the same public licensing facts are presented as a narrative, a table, or a simplified badge, how do research participants change their stated caution, comprehension, and willingness to contact a provider?
Why it matters. Ask Trust Hub’s parent role is routing and explanation. Presentation choices can create overconfidence even when the underlying record is honest.
Verticals. Cross-network
Potential output. Pre-registered experiment report. Marked as IRB-gated — not a live TrustHub test on production users.