Replication packages shared alongside published research sometimes contain personally identifiable information (PII) that was never intended for publication.
PII Checker helps researchers, data editors, and repositories scan replication packages for direct PII to help mitigate such unintended disclosure before publication.
PII Checker can be run locally on your own machine, or data editors and repositories can use our hosted service instead.
We used a version of this tool to find that 14.4% of 327 replication packages with author-collected data from 11 leading journals across economics, political science, and psychology contained direct identifiers like names, emails, or phone numbers. For details, see Brodeur and Valenta (2026).
This tool is designed to aid evaluation of already published replication packages, or packages that are about to be made publicly available.
Extract
Archives are unpacked, junk files cleaned out, and every data file in the package located, with duplicates identified along the way.
Screen columns
Every column across every file is screened against rule-based criteria to find the ones that need to be evaluated.
Check for PII
Each candidate column's name, label and values are evaluated by a (locally run) large language model for the presence of personally identifiable information
Report
A downloadable spreadsheet, one row per checked column, with its classification and the model's stated reasoning.
PII Checker is designed to evaluate tabular data. It does not search for PII in documents such as .pdf and .doc or in multimedia files. Currently the following formats are supported:
PII Checker produces one .xlsx report per package. Each evaluated column is assigned a classification and a short explanation of the model's reasoning, illustrated below.
See user guide for more details: User guide →
What type of PII is this tool designed to detect?
This tool is designed to detect direct identifiers, while columns might also be flagged as possible indirect identifiers, there is no holistic evaluation of the risk of reidentification of respondents.
What kind of data can this tool evaluate?
This tool only searches for PII in tabular data. Documents such as .pdf or .doc or multimedia files are not evaluated.
What file formats are supported?
Excel (.xls, .xlsx), plain text data (.csv, .tsv, .txt), Stata (.dta), R (.rds, .rdata), SPSS (.sav), SAS (.sas7bdat, .xpt), Matlab (.mat), and OpenDocument Spreadsheet (.ods) are supported.
Will the data be processed by third-party AI models?
No. To evaluate the presence of PII, the hosted version uses an open-source LLM (Gemma 4) that we run ourselves on our own servers hosted at Google Cloud - no third-party AI API is used, and your data isn't shared with anyone else. If you run PII Checker yourself, the default setup runs the same kind of open-source model locally via Ollama, so nothing leaves your machine either; a third-party API (Anthropic/OpenAI) is only ever used there if you deliberately opt into one for testing/benchmarking.
Does my data leave my computer, and how is it stored?
If you run PII Checker yourself with the default local Ollama setup, no data ever leaves your computer. If you use the hosted version instead, the uploaded package and any extracted files are deleted immediately once your job finishes, fails, or is cancelled; column values and model reasoning are stored as part of the evaluation output for one week, or until you delete the file. See the data, privacy & terms policy for the complete breakdown.
Who funds this tool?
This tool is supported by small amount of funding from the University of Ottawa, however we are actively seeking funding to support this project.
You can reach us using the following email: info [at] piichecker.org.