Methodology

What we do

The BlackCAT is a corporate accountability platform that helps Black people make informed decisions about where to work, shop, and bank. By turning fragmented public records of corporate harm into clear company profiles and actionable tools, it makes accountability easier to understand, pursue, and organize around.

Data Transparency and Trust

TheBlackCAT company record is designed to be transparent, explainable, and defensible. To achieve this, the proposed methodology moves away from opaque, model-based scoring in favor of a clear, hierarchy-based system grounded in verifiable public records.

Core Methodology

The system utilizes a weighted-average computation over severity-scored categories, presented to the community as tiered bands. This hybrid approach ensures:

  • Transparency: Every point in a score is traceable to a specific, saved public record.
  • Accountability: Raw records remain visible, allowing users to verify the data behind the record.
  • Neutrality: Model-based scoring (which risks obscuring logic and reproducing historical bias) is rejected to maintain public trust.

Architectural Scalability

To support the platform’s growth across multiple domains across a company (Employment launching in September 2026, followed by subsequent domains such as Donor Advisor Funds (DAF), Lending, DEI support, etc), the architecture is organized into a three-level hierarchy:

  • Domain (Top): Broad domain (e.g., Labor, DAF, Financial, DEI) that function as independent areas.
  • Category (Middle): Specific groupings (e.g., wage theft, lending discrimination).
  • Subcategory (Bottom): Granular data entries (e.g., child labor violations) that carry specific severity weights and evidence stages.

This structure allows TheBlackCAT to scale by adding new domains as branches within the existing hierarchy, avoiding the need to re-engineer the underlying scoring logic with each new data source.

Data Handling

How we solicit the data

To ensure each company record can be traced, we start by identifying public and private databases where data can be filtered by company and race based variables. We go broad in our search for data to identify public databases as well as private ones that are curated by trusted not-for-profit and private organizations. Each data source is clearly labeled and called out in our designs. Every community member can see where the records have been pulled from.

Data collection & classification

While the data scrapping is automated, every record is reviewed and approved by us before it is added to the final classification. Every record is saved so it can be easily referenced. We leverage Artificial Intelligence (AI) for the first round of classification. We conduct several statistical evaluations to ensure our categorization taxonomy is logical and fair. Our goal is to confirm that the way we grouped and labeled these violations accurately reflected real-world patterns. Additionally, we perform multiple diagnostics to validate the taxonomy:

  • Analyzing Frequency by Source: We compared how often each violation label appeared across the different data sources.
  • Checking Overlap (Co-occurrence): We examined how often violation labels appeared together in the same case. We found that the labels are not independent; for example, a single complaint often alleges multiple related issues, such as termination, retaliation, and harassment appearing simultaneously.
  • Identifying Underlying Structure: We employed rigorous statistical analysis to map the data landscape. Our findings confirm a robust, consistent structure, while revealing that a significant volume of multi-issue reports—where numerous grievances are filed simultaneously—shapes the overall data profile.
  • Clustering Labels: By grouping similar labels, we confirmed that our categorization wasn’t actually distinct, independent concepts but instead the data reflects one large cluster of conduct-based violations, with a few outliers.

Proprietary Scoring Process

Our methodology provides a data-driven view of corporate conduct by weighing several critical dimensions of an incident rather than simply counting cases. We use the following indicators to calculate the impact of corporate activity:

Severity

We assess the outcome of each record. Settled or adjudicated cases are weighed more heavily than unproven allegations.

Recency

To ensure our scores reflect current practices, we prioritize recent incidents within a 10-year rolling window, with weights decaying as cases age.

Breadth

We account for volume by tracking the number of cases per company, using a dampened multiplier to reflect the scale of affected individuals.

Recurrence

We distinguish between isolated clusters and systemic patterns by tracking how many distinct years a company appears with the same violation.

Irreversibility

We assign weights based on the lasting impact of a violation. Conduct that causes permanent harm (like termination) is prioritized over incidents that can be fully rectified (like administrative corrections).

Evidentiary Strength

Our scores are grounded in public records. We currently utilize verified data from federal courts and government agencies, with a framework to incorporate additional sources in the future.

Scope & identity

To ensure accuracy, we apply systemic modifiers to account for organizational-level practices and perform a two-tier cleaning process to aggregate data across subsidiaries and parent companies.

Testing for Stability

We designed our scoring system to be reliable, not arbitrary. To ensure our rankings are based on solid evidence rather than specific numerical choices, we subjected our methodology to rigorous stress testing. We randomly adjusted our weighting factors by 25% and recalculated the rankings 500 times. The results remained consistent, confirming that our system is stable. You can trust that our rankings are built on a solid foundation, not fragile calculations.

Each Company Harm Domain is distinct

While the first release focuses on one question: How have companies treated Black workers, each subsequent release will cover corporate political spending, lending discrimination, environmental harms, data-center locations, or the sale of surveillance technology. Those harm domains are distinct to each other. We will follow the same in-depth process that we used for Employment for each harm area.

Continuous Data Refresh

We will continue to update data from existing data sources as well as new data sources identified so the community always has the most up to date information for decision making. For example employment data feed will be updated periodically and we will be adding new data sources that meet our stringent criteria to ensure the community always has the most updated information.