BHRC Research Data Guide
Public aggregate summaries

9  BHRC Polygenic Score Database v3.0

Author: Lucas Toshio Ito - UNIFESP
lucas.toshio.ito@gmail.com

Supervisor: Marcos Leite Santoro – UNIFESP
santorogen@gmail.com

Genetics Coordinator: Síntia Belangero – UNIFESP
sinbelangero@gmail.com

Last updated: September 28, 2026


9.1 Overview

This database stores polygenic scores (PGS or PRS) for several phenotypes in the Brazilian High-Risk Cohort for Mental Health Conditions (BHRC). Out of the 2,511 probands in the study, 2,175 underwent genotyping and passed quality control (QC), so only these individuals can be used for genetic analysis.

We calculated the PGS for the participants using PRS-CS v1.1.0 (Ge et al., 2019) and PRS-CSx v1.1.0 (Ruan et al., 2022), based on summary statistics from prior genome-wide association studies (GWAS) for various phenotypes, primarily sourced from the Psychiatric Genomics Consortium (PGC). PRS-CSx was only applied for multi-ancestry GWAS phenotypes with ancestry-specific summary statistics available.

PRS-CS applies a continuous Bayesian shrinkage parameter to estimate the mean posterior effect size for a GWAS, incorporating an external 1000 Genomes European Linkage Disequilibrium (LD) panel. The global shrinkage parameter for PRS-CS was derived from the data using a fully Bayesian approach, with other parameters set to default. The PLINK v1.9 score function was used to calculate individual PGSs, which were then standardized in R to have a mean of 0 and a standard deviation of 1.

The PGS value is based solely on the individual’s genetic data and the summary statistics used. Therefore, regardless of the number of individuals analyzed, this value will remain consistent.

Publications that have used BHRC polygenic scores are listed on the BHRC polygenic score publications page.


9.2 Authorship Policy

All papers utilizing genetic data from the cohort should invite the Genetics Coordinator, Prof. Dra. Síntia Belangero – UNIFESP (sinbelangero@gmail.com), as a co-author. The Genetics Coordinator will determine who else should be included to assist with the analysis and be listed as co-authors on the paper.


9.3 Database Structure

All generated PGS were merged into a single file with the principal components in the last columns.

9.3.1 File format

  • PRSCS_”Group”_”Identifier”.csv: Main file with the PGS value of each individual and the principal components which should be used as covariates.
    • Group: Probands or Parents (Father/Mother) or ALL (both of them).
    • Identifier: identifier used for the file (ident, subjectid, probident, externalid, ext_genid).
    • PGS: Each phenotype has its own PGS column in the format “prscs/prscsx”_”Phenotype”_”Year”.
    • Principal components (PCs): Located at the end of the file and named PC1 through PC20.

9.4 Phenotypes Included

The full information about each variable is available in the PGS Data Dictionary. Files related to the following phenotypes are included in this folder:


9.5 Use with caution 1 – Principal Component Analysis (PCA)

These scores should be adjusted using the principal components (PCs) of the genotyping data in each subsequent analysis (e.g., as covariates in regression). Since the value of the principal components varies depending on the number of individuals included, you should not use the available PCs (N = 2175) if you’re working with a subsample.

We have calculated the first 20 PCs, but authors can choose how many to use as covariates. We recommend using 10 PCs, as it is more common and may be more easily accepted by reviewers.

To obtain a subsample-specific PC file, please email us with your request, copying your PI and the main contact in the cohort.


9.6 Use with caution 2 – Brazilian Admixture

Brazil is a highly admixed country, with ancestry proportions averaging around:

  • 65% European
  • 25% African
  • 10% Amerindigenous

These percentages vary widely across individuals and regions. Since most GWAS studies are predominantly based on European ancestry, prediction accuracy may be lower for our Brazilian sample. While representation of diverse ancestries in genetic studies has improved in recent years, it remains an important factor to consider when using and interpreting PGS results.


9.7 Use with caution 3 – Checking GWAS Updates

Please note that it is the responsibility of the authors to verify whether the GWAS used is the most up-to-date for the given phenotype. Using outdated data may lead to issues when publishing articles. If you require the PGS for a phenotype not included in the database or a more current PGS, please contact us.


9.8 Use with caution 4 – GWAS Sample Size

The sample size of the GWAS used matters significantly. If the PGS is based on a GWAS with only a few thousand participants and includes only a small number of genome-wide associated variants (e.g., some GWAS may have no significant hits below the p-value threshold), the PGS may not be very reliable. However, this may improve as future GWAS studies increase their sample sizes.


9.9 Use with caution 5 – Sample Relatedness

PC-Relate identified nine participant pairs (18 participants) above the internal relatedness threshold. Pair-level identifiers are maintained in a controlled private file and are not displayed on this public website.

Researchers should account for relatedness in the analysis plan, for example by defining an unrelated subset or using an analytical method that models relatedness. Contact the genetics team when pair-level information is required.


9.10 Use with caution 6 – Family-based Analysis

The PGS for the parents of the 2,511 probands have been added to this dataset, including:

  • 2144 mothers
  • 971 fathers

This forms:

  • 754 complete trios
  • 1110 mother–proband duos
  • 118 father–proband duos

Given the significant impact of relatedness in genetic studies, any project involving both probands and their parents must account for this factor to avoid bias in the results.

To maintain methodological rigor and ensure proper organization of the dataset, the PGS and PCs for parents are stored separately and are available upon request.


9.11 Updates from Version 2.0 and 3.0

  • New phenotypes have been added, and some existing ones have been updated based on the latest published GWAS.

  • The previous version of the PGS Database was based solely on genotyped variants. In this release, we have imputed the genetic data to include additional variants in the score calculation, enhancing prediction accuracy.

  • Initially, this database was created exclusively for probands. Now, PGS have been calculated for each phenotype in all parents with genotyped data as well. See Use with caution 6 – Family-based analysis.