Skip to content
ASKTRUST HUB

TrustHub Academic Research Program

Independent research using public regulatory and consumer-protection data

The TrustHub network organizes difficult public regulatory datasets across multiple consumer industries. This page is the professional foundation for selected research datasets, documentation, benchmarks, and project opportunities — not a claim that universities already use TrustHub.

Ask Trust Hub Standard

Last reviewed:

Program foundation · Preparing 2027 pilot projects

Why TrustHub data exists

Specialist TrustHub sites assemble structured public-source evidence so consumers — and, later, independent researchers — can inspect identity, licensing, and enforcement context without buying a ranking. Common ownership · Separated research and listing order · No paid placements. Evidence models differ by industry. Identical coverage across hubs is not claimed.

Ask Trust Hub is the parent research and standards layer. Directories and deep tools stay on specialist domains.

NetworkMethodologyData sourcesIndependenceTrust Center

Academic independence charter

These are program principles. They apply whether or not a university ever uses the materials.

  1. 01

    Unfavorable findings are allowed

    Researchers are free to reach conclusions that are unfavorable to TrustHub, to a specialist hub, or to businesses represented in public records.

  2. 02

    No requirement of favorable conclusions

    TrustHub does not require, incentivize, or condition access on favorable conclusions about the network, its methodology, or any listed entity.

  3. 03

    Factual correction, not suppression

    TrustHub may correct factual errors about an underlying dataset or explain source limitations. It must not suppress findings because they are unfavorable.

  4. 04

    Methodology critique is a feature

    Findings that expose weaknesses in TrustHub methodology should be treated as opportunities to improve the system, not as reputational incidents to bury.

  5. 05

    No consumer PII

    The academic program does not provide consumer personally identifiable information. Releases must be designed so that consumer PII is not included.

  6. 06

    Limitations must be documented

    Dataset limitations, missing evidence, uneven state coverage, and overwrite risk in source systems must be documented alongside any future release.

  7. 07

    Original sources remain the authority

    Public regulatory data should remain attributable to its original source. TrustHub organization of records is not a substitute for the underlying registry.

  8. 08

    Researchers retain analytical independence

    Academic researchers retain independence over their questions, methods, interpretation, and publication decisions.

  9. 09

    Review is not editorial approval

    Future collaboration agreements may allow a limited factual or legal review period before publication. Review is never approval based on favorability of findings.

  10. 10

    Documented methodology changes

    When independent research leads to a meaningful change in TrustHub methodology, TrustHub should publicly document that change when appropriate.

Year-one research focus

The initial program stays narrow. Two primary tracks only. Other disciplines are listed as possible later expansion — not as open offerings.

Primary track

Data science / AI

Entity resolution, missing-data bias, calibration, longitudinal snapshots, and measurement of matching systems against public business identity records.

Primary track

Public policy / consumer protection

How regulatory information is published, overwritten, or withheld across states and agencies — and whether that structure helps or hinders consumers.

Future expansion areas: Gerontology, Consumer law, Finance, Insurance, Construction management. Not launched.

Engagement model

Tier 0

Planned

Open research

Versioned public research datasets, documentation, methodology notes, and citation information where legally and operationally appropriate. Nothing in this catalog is a public download yet.

Tier 1

Planned

Classroom research

Turnkey project materials suitable for classroom assignments, independent study, or research-methods courses once datasets and documentation are released.

Tier 2

Planned

Capstone / practicum

A defined real-world research problem, documented dataset, and limited TrustHub support for a student team. Coordinators — not cold outreach to individual faculty — are the preferred first audience.

Tier 3

Future · not currently open

Funded research

A future possibility only. TrustHub is not soliciting funded proposals, creating payment systems, or promising grants in this foundation.

TrustHub is not promising funding and is not soliciting funded proposals.

Research dataset registry (preview)

No academic datasets are published yet. Entries below are prospective research areas aligned to evidence each specialist hub already uses. There are no download links and no DOIs.

  • Move Trust HubPlanned — not released

    Mover regulatory identity and authority (prospective)

    Organized public FMCSA/SAFER-oriented identity for household-goods movers: operating authority context, broker/carrier classification where published, and enforcement or fitness history that Move Trust Hub already treats as research evidence.

    Sources: FMCSA · SAFER · Access: internal research · DOI: none assigned · Download: not available

  • SeniorTrustHubDocumentation in progress

    Nursing-facility CMS records and ownership context (prospective)

    CMS-oriented nursing-facility identity, ownership relationships where CMS publishes them, and inspection or enforcement history that SeniorTrustHub already treats as government-sourced research — not placement inventory.

    Sources: CMS · Supported state regulators · Access: internal research · DOI: none assigned · Download: not available

  • Contractor Trust HubPlanned — not released

    State contractor-license records and multi-state identity (prospective)

    Official-board contractor license and registration identity with state-specific depth, including disciplinary evidence where a board already publishes it on Contractor Trust Hub.

    Sources: State contractor licensing boards · Access: internal research · DOI: none assigned · Download: not available

  • InvestorTrustHubPlanned — not released

    SEC / IARD investment-adviser firm records (prospective)

    Firm-level SEC/IARD registration identity as used on InvestorTrustHub. This is not a FINRA BrokerCheck people product, stock-advice dataset, or complete advisor-profile dump.

    Sources: SEC · IARD · Access: internal research · DOI: none assigned · Download: not available

  • Insurance Trust HubPlanned — not released

    Insurance agency / producer regulatory identity (prospective)

    State DOI / NAIC-oriented identity for insurance agencies and producers where Insurance Trust Hub already organizes public license context. Coverage is not a single national producer database.

    Sources: State Departments of Insurance · NAIC · Access: internal research · DOI: none assigned · Download: not available

  • Lender Trust HubPlanned — not released

    Mortgage-company NMLS-oriented public identity (prospective)

    Company-level NMLS Consumer Access identity and related public records (CFPB complaint patterns, FDIC/bank identity where relevant) as Lender Trust Hub already uses them for research — not a lead list and not a universal Trust Score dump.

    Sources: NMLS Consumer Access · CFPB · FDIC / public banking records · Access: internal research · DOI: none assigned · Download: not available

  • Cross-networkDocumentation in progress

    TrustHub Entity Resolution Benchmark (unlabeled candidates)

    Internal unlabeled candidate pairs for later dual human review of difficult Move (FMCSA) and Contractor (state license) business-identity questions. Similarity and matcher confidence are not ground truth.

    Sources: FMCSA · State contractor licensing boards · Access: internal research · DOI: none assigned · Download: not available

Versioned, longitudinal data

Future academic releases should be immutable, date-stamped, reproducible, source-attributed, documented, and retained historically rather than overwritten. Public registries often replace prior state; a snapshot can become more useful with time for that reason.

This page states the contract. It does not run a production snapshot pipeline.

Citation and DOI readiness

TrustHub Research Data. [Dataset Title]. Version [X]. [Release date]. DOI: [when assigned].

  • Do not invent a DOI. Leave the DOI field blank until a repository assigns one.
  • Version and release date refer to an immutable snapshot, not a live production extract.
  • Cite the original regulator (FMCSA, NMLS, CMS, SEC, state boards, DOI/NAIC) as the source of the underlying records.
  • TrustHub citation covers the organized snapshot, schema, and documentation — not ownership of the public records.

Possible later archives (not registered): Zenodo, Harvard Dataverse.

TrustHub Entity Resolution Benchmark

Documentation and internal unlabeled candidates (Academic 001C.1). No public ground-truth file, no claimed precision/recall, no production matching dump.

Future example classes

  • Legal name versus DBA
  • Same entity, different location strings
  • Company rename / successor-predecessor
  • Duplicated regulatory records
  • Corporate-family members that should not collapse
  • Ambiguous near-matches
  • False-positive traps (similar names, unrelated firms)

Future metrics (not computed)

  • Precision
  • Recall
  • False-positive rate
  • False-negative rate
  • Confidence calibration

Turnkey project library

12 research questions for a future classroom or capstone library. They are not assignments currently in the field. Projects involving human participants are marked for likely IRB review.

Public policycapstone · advanced

Can regulatory enforcement predict subsequent consumer complaints?

Question. After a public enforcement or disciplinary event appears in a regulator’s records, do later consumer-complaint patterns change in a measurable way — and in which verticals is the relationship weakest?

Why it matters. Consumers and journalists often treat a single enforcement record as a lasting risk signal. Measuring lag, attenuation, and missing-data bias would test that assumption without implying TrustHub scores.

Verticals. Move Trust Hub, Lender Trust Hub, Contractor Trust Hub

Potential output. Working paper plus a replication notebook using a future versioned snapshot — not a TrustHub ranking product.

Public policymasters · substantial

Geographic variation in regulatory enforcement intensity

Question. How much of the observed difference in published enforcement counts across U.S. states is explained by industry size versus publication practice versus genuine enforcement intensity?

Why it matters. State boards and DOIs do not publish equally. Treating raw counts as “risk maps” can punish transparency rather than misconduct.

Verticals. Contractor Trust Hub, Insurance Trust Hub, SeniorTrustHub

Potential output. State-comparison memo with a documented incompleteness appendix.

Data science / AIcapstone · advanced

Measuring entity-resolution precision and recall across public business registries

Question. Using a future hand-reviewed TrustHub Entity Resolution Benchmark, what precision, recall, and false-positive rate would a simple name-and-place matcher achieve versus a more conservative exact-identifier matcher?

Why it matters. Public registries duplicate, rename, and DBA-shift. Over-merging unrelated firms is a consumer-protection failure, not a convenience.

Verticals. Cross-network

Potential output. Benchmark card and error taxonomy. No claimed TrustHub performance in Academic 001A.

Data science / AIcapstone · advanced

Phoenix-company detection in public records

Question. Can public regulatory records identify operators that dissolve or abandon an identity and later reappear under a new business name, address cluster, or officer set — without claiming a legal finding of fraud?

Why it matters. Household-goods moving and some construction markets have a documented public-interest problem of successive identities. Measurement should stay evidence-based and cautious.

Verticals. Move Trust Hub, Contractor Trust Hub

Potential output. Method paper emphasizing uncertainty, with withheld names unless counsel clears a public version.

Data science / AImasters · substantial

Calibration audit of TrustHub-derived research signals (where they exist)

Question. Where a specialist hub publishes an explicit research score or composite signal, how well do those values calibrate to later independently observed outcomes — and where a hub does not publish a proprietary Trust Score, what happens if researchers invent one anyway?

Why it matters. Not every hub uses a proprietary Trust Score. Forcing a universal score across moving, lending, insurance, contractors, senior care, and investment firms would misstate the methodology.

Verticals. Cross-network

Potential output. Hub-by-hub methods note: scored vs unscored evidence models.

Public policyundergraduate · moderate

Regulatory transparency across U.S. states: licensing and disciplinary records

Question. How accessible are licensing lookups and disciplinary records for contractors and insurance producers across states — by search UX, bulk data, and historical retention?

Why it matters. Consumer protection depends on whether the public can actually retrieve the record, not only whether a statute exists.

Verticals. Contractor Trust Hub, Insurance Trust Hub

Potential output. Open codebook and state scorecard of publication practice — not of “bad actors.”

Data science / AImasters · substantial

Missing-data bias in consumer regulatory datasets

Question. When a field, county, or state is systematically missing from public extracts, which research conclusions flip if missingness is treated as “clean” versus “unknown”?

Why it matters. TrustHub’s public stance is that missing evidence does not mean a clean record. That claim is testable.

Verticals. Cross-network

Potential output. Methods appendix suitable for later dataset dictionaries.

Data science / AIundergraduate · moderate

Business-name normalization and DBA matching

Question. How much of apparent “duplicate company” volume in public registries is explained by punctuation, legal-suffix, and DBA variation versus genuinely distinct entities?

Why it matters. Name cleaning is the cheapest matcher and the easiest way to create false merges.

Verticals. Move Trust Hub, Lender Trust Hub, InvestorTrustHub

Potential output. Normalization spec that a later academic snapshot can cite.

Public policycapstone · advanced

Ownership networks and nursing-facility regulatory outcomes

Question. Where CMS publishes ownership or chain relationships, how do inspection or enforcement patterns cluster — and what cannot be claimed because ownership files are incomplete or lagged?

Why it matters. Nursing-facility ownership is a live public-policy subject. Research must stay inside CMS (and supported state) evidence, not placement marketing.

Verticals. SeniorTrustHub

Potential output. Policy brief with a limitations box on CMS lag and unpublished state actions.

Data science / AIindependent study · substantial

Measuring regulatory-data freshness and historical information loss

Question. When public systems overwrite prior license status, how much historical information is lost between two calendar dates — and what research designs become impossible without immutable snapshots?

Why it matters. A historical snapshot can become more valuable with time precisely because regulators overwrite. This project motivates the versioning standard without running a production snapshot job.

Verticals. Cross-network

Potential output. Protocol for longitudinal academic releases (see dataset-release standard).

Data science / AImasters · advanced

Algorithmic confidence calibration for business identity matching

Question. If a matcher emits a confidence score, does predicted confidence equal empirical match frequency across bins — or are high-confidence errors concentrated in DBA and successor cases?

Why it matters. Uncalibrated “95% match” labels are how false merges get into consumer-facing products.

Verticals. Cross-network

Potential output. Calibration report template for later Academic 001B+ benchmark files.

Public policycapstone · substantial · IRB likely

Information presentation and consumer decision quality

Question. When the same public licensing facts are presented as a narrative, a table, or a simplified badge, how do research participants change their stated caution, comprehension, and willingness to contact a provider?

Why it matters. Ask Trust Hub’s parent role is routing and explanation. Presentation choices can create overconfidence even when the underlying record is honest.

Verticals. Cross-network

Potential output. Pre-registered experiment report. Marked as IRB-gated — not a live TrustHub test on production users.

Preferred audiences (later)

Outreach is not part of this foundation. When it begins, it should not start as cold-emailing individual professors.

  1. University capstone and practicum coordinators
  2. Interdisciplinary centers and research institutes
  3. Selected data science / AI faculty (after coordinator pathways exist)
  4. Selected public policy / consumer-protection faculty (after coordinator pathways exist)
  5. Later academic disciplines listed as expansion tracks

Deferred: University datathons; Data competitions; Guest lectures; Classroom research projects at scale; Cold-emailing individual professors as the primary strategy.

How we will judge success

Raw partnership count is not the primary metric.

  • Completed citable research outputs. The program exists to produce independent analysis, not partnership announcements.
  • Academic dataset citations. Citation of versioned releases is a better signal than logo placement.
  • DOI citations (when DOIs exist). Formal identifiers support replication and archival credit. None are assigned yet.
  • Dataset downloads where appropriate. Use only after an open release exists; do not treat unpublished catalogs as usage.
  • Returning faculty and programs. Repeat use indicates research utility, not a one-time courtesy listing.
  • Methodology improvements triggered by outside research. Unfavorable findings that change TrustHub practice are a success, not a failure.
  • Replicated analyses. Independent reproduction is stronger than a single flattering paper.
  • Student recruiting outcomes. Secondary: whether graduates who used the materials become stronger researchers.
  • High-quality external references. Researchers, regulators, or journalists citing the work for its substance.

Discuss an academic research project

Use the existing Ask Trust Hub contact path. Include “academic research” in the subject. There is no researcher portal, mailing list, or student-data form on this page.