PRS-CSx
Cross-population continuous-shrinkage scoring
PRS-CSx jointly models GWAS summary statistics from multiple populations using coupled continuous-shrinkage priors. It produces population-specific posterior effects and can also return meta-analyzed effects.
Prefer a reusable file? Download the PRS-CSx shell template, edit its inputs, and begin with the chromosome 22 test.
Install and check
git clone https://github.com/getian107/PRScsx.git
python -m pip install numpy scipy h5py
python PRScsx/PRScsx.py --helpPRS-CSx uses the same population-specific LD blocks as PRS-CS, plus a shared SNP-information file placed in the parent reference directory. See the official PRS-CSx repository for current downloads.
For PRS-CSx, --ref_dir points to the parent directory containing the SNP-information file and the required population-specific LD folders. This differs from the population-specific --ref_dir used by PRS-CS.
Define the inputs
PRSCSX_DIR=/path/to/PRScsx
REF_DIR=/path/to/reference
TARGET_PREFIX=/path/to/target/target_genotypes
EUR_SUMSTATS=/path/to/gwas/eur_sumstats.txt
EAS_SUMSTATS=/path/to/gwas/eas_sumstats.txt
OUTPUT_DIR=/path/to/results
OUTPUT_NAME=prscsx_test
mkdir -p "${OUTPUT_DIR}"
export MKL_NUM_THREADS=1
export NUMEXPR_NUM_THREADS=1
export OMP_NUM_THREADS=1Keep the inputs aligned
These three arguments are positional and must use the same population order:
--sst_file=eur_sumstats.txt,eas_sumstats.txt
--n_gwas=200000,100000
--pop=EUR,EAS
List the GWAS files, sample sizes, and population labels in the same population order. In the example above, eur_sumstats.txt, 200000, and EUR describe the first population; eas_sumstats.txt, 100000, and EAS describe the second.
Test chromosome 22
python "${PRSCSX_DIR}/PRScsx.py" \
--ref_dir="${REF_DIR}" \
--bim_prefix="${TARGET_PREFIX}" \
--sst_file="${EUR_SUMSTATS},${EAS_SUMSTATS}" \
--n_gwas=200000,100000 \
--pop=EUR,EAS \
--chrom=22 \
--meta=True \
--out_dir="${OUTPUT_DIR}" \
--out_name="${OUTPUT_NAME}"With --meta=True and no fixed --phi, PRS-CSx produces the auto/meta solution described in the official documentation. This is useful when a separate validation dataset is unavailable.
Run the autosomes
for CHR in $(seq 1 22); do
python "${PRSCSX_DIR}/PRScsx.py" \
--ref_dir="${REF_DIR}" \
--bim_prefix="${TARGET_PREFIX}" \
--sst_file="${EUR_SUMSTATS},${EAS_SUMSTATS}" \
--n_gwas=200000,100000 \
--pop=EUR,EAS \
--chrom="${CHR}" \
--meta=True \
--out_dir="${OUTPUT_DIR}" \
--out_name="${OUTPUT_NAME}"
doneOn a cluster, submit chromosomes as separate jobs instead of running the loop sequentially.
Inspect and combine outputs
PRS-CSx writes one set of population-specific posterior effects per input GWAS. With --meta=True, it also writes meta-analyzed effects.
find "${OUTPUT_DIR}" -maxdepth 1 -name "${OUTPUT_NAME}*chr*.txt" -printChoose the intended population-specific or meta-analyzed files, then concatenate only that matched series:
cat /path/to/the/matched_chr*.txt \
> "${OUTPUT_DIR}/${OUTPUT_NAME}_pst_eff_all.txt"Avoid a broad wildcard that mixes population-specific and meta-analyzed results.
Calculate a score
plink \
--bfile "${TARGET_PREFIX}" \
--score "${OUTPUT_DIR}/${OUTPUT_NAME}_pst_eff_all.txt" 2 4 6 sum center \
--out "${OUTPUT_DIR}/${OUTPUT_NAME}_score"Choosing the final PRS-CSx strategy
The official documentation describes two main approaches:
- Validation-based linear combination: calculate one score per discovery population, standardize them, and learn their weights in an independent validation dataset.
- Auto/meta: use
--meta=Truewithout tuning a fixedphi, producing a combined score when a validation dataset is not available.
The validation-based combination may provide better performance for a specific target population, but its weights and hyperparameters must be evaluated in an independent test dataset.
Evaluate in R
Use the same workflow shown on the PRS-CS page: merge scores, phenotype data, and ancestry PCs by participant identifiers; standardize the score; and compare a covariate-only model with a model that adds the PRS.
Cross-population prediction does not remove ancestry-related limitations. Document the discovery populations, LD references, target population, score-combination strategy, and validation procedure.