plink
- Version:
1.9b_6.21, 2.00a2.3, 2.00a3.6
- Category:
bio
- Cluster:
Vali
Description
PLINK is a widely-used open-source toolset for whole-genome association and population-based linkage analyses.
Key capabilities:
Genome-wide association studies (GWAS)
Population genetics and linkage analysis
Data management and format conversion
Variant filtering and quality control
Supports multi-threaded execution for performance
Note
PLINK provides two main versions: plink (1.9) and plink2 (2.0). Make sure to use the correct command based on the module loaded
Documentation
> plink2 --help
PLINK v2.00a2.3 AVX2 (24 Jan 2020) www.cog-genomics.org/plink/2.0/
(C) 2005-2020 Shaun Purcell, Christopher Chang GNU General Public License v3
plink2 <input flag(s)...> [command flag(s)...] [other flag(s)...]
plink2 --help [flag name(s)...]
Most PLINK runs require exactly one main input fileset. The following flags
are available for defining its form and location:
--pfile <prefix> ['vzs'] : Specify .pgen + .pvar[.zst] + .psam prefix.
--pgen <filename> : Specify full name of .pgen/.bed file.
--pvar <filename> : Specify full name of .pvar/.bim file.
--psam <filename> : Specify full name of .psam/.fam file.
--bfile <prefix> ['vzs'] : Specify .bed + .bim[.zst] + .fam prefix.
--bpfile <prefix> ['vzs'] : Specify .pgen + .bim[.zst] + .fam prefix.
--keep-autoconv : When importing non-PLINK-binary data, don't delete
autogenerated binary fileset at end of run.
--no-fid : .fam file does not contain column 1 (family ID).
--no-parents : .fam file does not contain columns 3-4 (parents).
--no-sex : .fam file does not contain column 5 (sex).
--vcf <filename> ['dosage='<field>]
--bcf <filename> ['dosage='<field>] (not implemented yet) :
Specify full name of .vcf{|.gz|.zst} or BCF2 file to import.
* These can be used with --psam/--fam.
* By default, dosage information is not imported. To import the GP field
(must be VCFv4.3-style 0..1, one probability per possible genotype), add
'dosage=GP' (or 'dosage=GP-force', see below). To import Minimac3-style
DS+HDS phased dosage, add 'dosage=HDS'. 'dosage=DS' (or anything else
for now) causes the named field to be interpreted as a Minimac3-style
dosage.
Note that, in the dosage=GP case, PLINK 2 collapses the probabilities
down to dosages; you cannot use PLINK 2 to losslessly convert VCF
FORMAT:GP data to e.g. BGEN format. To make this more obvious, PLINK 2
now errors out when dosage=GP is used on a file with a FORMAT:DS header
line and --import-dosage-certainty wasn't specified, since dosage=DS
extracts the same information more quickly in this situation. You can
suppress this error with 'dosage=GP-force'.
In all of these cases, hardcalls are regenerated from scratch from the
dosages. As a consequence, variants with no GT field can now be
imported; they will be assumed to contain only diploid calls when HDS is
also absent.
--data <filename prefix> [REF/ALT mode] ['gzs']
--bgen <filename> [REF/ALT mode] ['snpid-chr']
--gen <filename> [REF/ALT mode]
--sample <filename> :
Specify an Oxford-format dataset to import. --data specifies a .gen[.zst]
+ .sample pair, while --bgen specifies a BGEN v1.1+ file.
* If a BGEN v1.2+ file contains sample IDs, it may be imported without a
companion .sample file.
* With 'snpid-chr', chromosome codes are read from the 'SNP ID' field
instead of the usual chromosome field.
* The following REF/ALT modes are supported:
'ref-first': The first allele for each variant is REF.
'ref-last': The last allele for each variant is REF.
'ref-unknown' (default): The last allele for each variant is treated as
provisional-REF.
This parameter will be required instead of optional in alpha 3.
--haps <filename> [{ref-first | ref-last}]
--legend <filename> <chr code> :
Specify .haps [+ .legend] file(s) to import.
* When --legend is specified, it's assumed that the --haps file doesn't
contain header columns.
* On chrX, the second male column may contain dummy '-' entries. (However,
PLINK 2 currently cannot handle omitted male columns.)
* If not used with --sample, new sample IDs are of the form 'per#/per#'.
Output files have names of the form 'plink2.<extension>' by default. You can
change the 'plink2' prefix with
--out <prefix> : Specify prefix for output files.
Examples/Usage
Load the PLINK module:
$ module load PLINK/2.00a2.3-GCC-10.3.0
Unload the module:
$ module unload PLINK/2.00a2.3-GCC-10.3.0
Start MATLAB with full desktop environment:
$ plink2
Serial Jobs
Here is an example job for PLINK version 1.9 running on 4 cores and 8GB of memory:
#!/bin/bash
#SBATCH -n 4 # (or --ntasks=4) Request 4 cores
#SBATCH --mem-per-cpu=2G # Request 2GB RAM per core
#SBATCH -t 1:0:0 # Request 1 hour runtime
module load plink/1.9-beta6.27-gcc-12.2.0
plink --threads ${SLURM_NTASKS} \
--silent \
--dummy 387 1112 0.03 scalar-pheno \
--out dummy1
plink --threads ${SLURM_NTASKS} \
--bfile dummy1 \
--recode \
--out dummy1.new_recode
Here is an example job for PLINK version 2.0 running on 4 cores and 8GB of memory:
#!/bin/bash
#SBATCH -n 4 # (or --ntasks=4) Request 4 cores
#SBATCH --mem-per-cpu=2G # Request 2GB RAM per core
#SBATCH -t 1:0:0 # Request 1 hour runtime
module load plink
plink2 --threads ${SLURM_NTASKS} \
--dummy 387 1112 0.03 scalar-pheno \
--out dummy1
plink2 --threads ${SLURM_NTASKS} \
--pfile dummy1 \
--export vcf \
--freq \
--keep-autoconv \
--out dummy1.new_recode
Installation
PLINK is pre-installed on the cluster and available via modules. Load the appropriate module and use directly.