Many bioinformatics tools exist to annotate reconstructed transcripts but are too complex to use for non–bioinformaticians. We have developed a platform to automatically annotate large sets of transcriptomic sequences. Annotations can be inferred based on their similarities with reference annotated sequences. Annotations can also be inferred based on structure and domains of protein sequences. Genotate is available on Genotate Website.
We provide an interactive and intuitive web platform, named Genotate, which allow to search and visualize this identified annotations. Genotate web platform can be installed on a server, follow the instruction on Github.
Genotate unifies ORF identification, similarity annotation and functional annotation in an automatic annotation platform. Genotate provides to biologist the possibility to annotate transcript sequence and entire transcriptome.
Genotate is able to identify for each reconstructed transcript all the ORF. To search for a peptidic functional domain it is necessary to know the encoded protein sequence.
For each transcript to analyze, Genotate first check all transcript with CPAT to measure their coding potential. CPAT use by default the longest ORF which is kept if the coding potential is superior the selected threshold. ORFs are then translated to obtain the associated protein sequences. The user can also chose to disable CPAT and use a custom detection of all possible ORFs based on parameters selected in the ORF panel. The start and stop codons (which initiate and end the ORFs) can be specified by users. By default, start codon is set to 'ATG' and stop codons to 'TAG, TGA, and TAA'. ORFs with a length lower than a threshold can be filtered to avoid interpretations of sequences with no biological meaning. Inner ORFs (which consist of nested ORF sequences) can also be identified as well as outside ORF (which consist of ORFs lacking either the start or stop codon). By default, the transcript without any coding ORF are annotated as a non-coding RNA; and ORFs are also identified on the reverse complemented transcript sequence.
In detail, the protein associated to a transcript can be obtained from multiples ORF encoded on the transcript. A frame is composed of nucleotide triplets called codon. The transcript sequence is divided into three frames, with a shift of one base on the sequence strand. The transcript sequence can also be reversed, and the nucleic base complemented to obtain the complementary sequence. An Open Reading Frame begins with a codon start and ends with a codon stop. A codon can be translated to an amino acid or end of translation signal. A codon encoding the beginning of the translation, such as 'ATG', is called codon start. A codon encoding the end of the translation, such as 'TAG, TGA, TAA', is called codon stop. A protein is obtained from the translated sequence of an Open Reading Frame. Moreover, due to sequencing errors the start or the stop codon can be truncated and not included in the input transcript sequence.
The following figure represents the diversity of ORF. An ORF can begin at every start codon which has a stop codon in the same frame.
ORF 2 and 6 are complete ORF which includes the inner ORF 1, 4 and 5. The ORF 3 and 7 are incomplete and characterized as outside ORF.
Genotate annotate sequences based on the similarity to other reference sequences.
The similarity annotations are computed based on any set of reference sequences specified by the user. The similarity annotations are identified at both the nucleic and peptidic levels. Sequences similarities are identified with the BLAST algorithm. Similarity results can be filtered based on the identity percentage. Similarity results can also be filtered based on the query sequence cover percentage and subject sequence cover percentage. The descriptions associated with the matched sequences are available in the similarity results computed by Genotate.
| Name | Information | Link |
| Ensembl | Genome, transcriptome(cds, cdna, ncrna) and proteome for a large number of species | Link |
| NCBI | Genome, transcriptome(cds, cdna, ncrna) and proteome for a large number of species | Link |
| NONCODE | database dedicated to non-coding RNAs (excluding tRNAs and rRNAs), by species | Link |
| Uniprot | UniProtKB/TrEMBL contains high quality computationally analyzed records that are enriched with automatic annotation and classification. Swissprot contains high quality manually annotated and non–redundant protein sequence database. | Link |
Genotate is able to identify functional domains based on multiples algorithm.
The functional annotations are computed based on a large set of publicly available computational tools and databases. Especially, we use the InterproScan software to find protein functional domains on the reconstructed transcript. InterproScan unifies proteins functional domains from different databases such as PFAM, SUPERFAMILY, and PANTHER. The PFAM database contains non redundant conserved proteomic functional domains found on various species. The SUPERFAMILY database provides information about the 3D structure information based on protein similarities. The PANTHER database provides ontology information with the molecular function and the biological process.
| Name | Information | Institute | Link |
| Interproscan | Standalone which unify protein analysis tool and databases | EMBL–EBI in Hinxton, The Wellcome Genome Campus | Link |
| Cdd | Search Conserved Domains and Protein Classification | The National Center for Biotechnology Information (NCBI) is part of the United States National Library of Medicine (NLM), a branch of the National Institutes of Health. | Link |
| Coils | Predicts coiled–coil conformation | SIB Swiss Institute of Bioinformatics | Link |
| Gene3d | Search CATH domain families from PDB structures | UCL Department of Biochemical Engineering. University College, London, UK. | Link |
| Hamap | classification and annotation system of protein sequences | SIB Swiss Institute of Bioinformatics | Link |
| Mobidblite | predictions of long intrinsically disordered regions | Department of Biomedical Sciences, University of Padua | Link |
| Panther | Gene ontology classification system | University of Southern California, CA, US. | Link |
| Pfam | Search protein families from Pfam database | EMBL European Bioinformatics Institute | Link |
| Pirsf | Search against fully curated PIRSF families with HMM models | Georgetown University Medical Center, University of Delaware | Link |
| Prints | Search protein fingerprints | University of Manchester, UK. | Link |
| Prodom | Search protein domain | PRABI Villeurbanne, France. | Link |
| Prosite | Search protein families and domains | Swiss Institute of Bioinformatics (SIB), Geneva, Switzerland. | Link |
| Sfld | Search enzymes classification in the Structure–Function Linkage Database | UC San Francisco, Babbitt Lab, SFLD Team | Link |
| Smart | Simple Modular Architecture Research Tool | EMBL, Heidelberg, Germany. | Link |
| Superfamily | structural and functional annotation for proteins and genomes | University of Bristol, UK. | Link |
| Tigrfam | identify functionally related proteins based on sequence homology | J. Craig Venter Institute, Rockville, MD, US. | Link |
The functional annotations can also be computed based on various prediction tools, such as TMHMM, SIGNALP, PROP. TMHMM predicts of transmembrane domains, which fundamentally rule all the membrane biochemical processes, with the hidden Markov models. SIGNALP predicts the secretory signal peptide, a ubiquitous signal that targets for translocation across the membrane, based on neural network. PROP predicts arginine and lysine propeptide, which characterize inactive peptides precursors which undergo post–translational processing to become biologically active polypeptides. The parameters of each tool available through Genotate can be specified, and the functional annotations can be filtered based on the evalue and other criteria specific to each tool.
| Name | Information | Institute | Link |
| BEPIPRED | Predict the location of linear B cell epitopes | Center for Biological Sequence Analysis, BioCentrum-DTU, Building 208, Technical University of Denmark | Link |
| MHCI | MHC I from IEDB database determine the ability to bind to a specific MHC class I molecule | Division of Vaccine Discovery, La Jolla Institute for Allergy and Immunology | Link |
| MHCII | MHC II from IEDB database predict MHC Class II epitopes | Division of Vaccine Discovery, La Jolla Institute for Allergy and Immunology | Link |
| NETCGLYC | NetCGlyc produces neural network predictions of C-mannosylation sites in mammalian proteins. | Department of Medical Biochemistry and Biophysics, Karolinska Institutet, SE-171 77 Stockholm, Sweden and Stockholm Bioinformatics Center | Link |
| NETNGLYC | NetNglyc predicts N-Glycosylation sites in human proteins using artificial neural networks | Center for Biological Sequence Analysis, The Technical University of Denmark, Lyngby, Denmark | Link |
| Prop | predicts arginine and lysine propeptide cleavage sites | Center for Biological Sequence Analysis at the Technical University of Denmark | Link |
| rnammer | Annotates ribosomal RNA genes | Centre for Molecular Biology and Neuroscience and Institute of Medical Microbiology, University of Oslo | Link |
| Signalp | predicts the presence and location of signal peptide cleavage sites in amino acid sequences | Center for Biological Sequence Analysis at the Technical University of Denmark | Link |
| Tmhmm | predicts of transmembrane helices in proteins | Center for Biological Sequence Analysis at the Technical University of Denmark | Link |
| tRNAscan | Predicts transfer RNA genes | Biomolecular Engineering, University of California Santa Cruz | Link |
Java is used to run Genotate. For similarity annotation, NCBI BLAST is used and need to be installed. For functional domain research, the tools InterproScan, TMHMM, SignalP, ProP and others need to be installed. After installation check IN THE EXECUTABLE SCRIPTS the location of tch, perl, python, libraries and temporary folders for all tools.
Java is used to launch Genotate and can be downloaded at java.com. Please install java in services/java/bin/java.
CPAT is used for coding-potential assessment, and need to be installed. CPAT can be downloaded at link.
wget https://sourceforge.net/projects/rna-cpat/files/v2.0.0/CPAT-2.0.0.tar.gz
pip3 install CPAT-VERSION.tar.gz
mv CPAT-VERSION CPAT
NCBI BLAST is used for similarity annotation, and need to be installed. BLAST can be downloaded at NCBI. BLAST is used for the annotation by homology. BLAST generate a database from a fasta nucleic and proteic file. Query sequence can then be aligned against the databases to search for similar sequences.
wget ftp://ftp.ncbi.nlm.nih.gov/blast/executables/blast+/LATEST/ncbi–blast–2.6.0+–x64–linux.tar.gz
tar –pxvzf ncbi–blast–2.6.0+–x64–linux.tar.gz
mv ncbi–blast–2.6.0+ blast
InterproScan is used for functional domain research, and need to be installed. InterProScan allows sequences to be scanned against functional domains, provided by several different databases. InterProScan can be downloaded at interproscan github.
wget ftp://ftp.ebi.ac.uk/pub/software/unix/iprscan/5/5.22–61.0/interproscan–5.22–61.0–64–bit.tar.gz
wget ftp://ftp.ebi.ac.uk/pub/software/unix/iprscan/5/5.22–61.0/interproscan–5.22–61.0–64–bit.tar.gz.md5
md5sum –c interproscan–5.22–61.0–64–bit.tar.gz.md5
tar –pxvzf interproscan–5.22–61.0–64–bit.tar.gz
mv interproscan–5.22–61.0 interproscan
To install Panther, please download the latest Panther data files (~ 12 GB). The data file need to be extracted into the [InterProScan5 home]/data/ directory. You can check the MD5 checksum on the file after you have downloaded it.
wget ftp://ftp.ebi.ac.uk/pub/software/unix/iprscan/5/data/panther–data–11.1.tar.gz
wget ftp://ftp.ebi.ac.uk/pub/software/unix/iprscan/5/data/panther–data–11.1.tar.gz.md5
md5sum –c panther–data–11.1.tar.gz.md5
tar –pxvzf panther–data–11.1.tar.gz
The full path to JAVA must be set in the environment or directly in the script interproscan.sh
JAVA=/var/www/genotate/services/java/bin/java
Edit interproscan.properties to change the number of parallel jobs allowed for interproscan and interproscan workers. At least one worker is required, which can launch other workers.
number.of.embedded.workers=1
maxnumber.of.embedded.workers=32
worker.number.of.embedded.workers=4
worker.maxnumber.of.embedded.workers=4
PROP predicts arginine and lysine propeptide cleavage sites, which characterize inactive peptides precursors. The precursors undergo post–translational processing to become biologically active polypeptides. ProP can be downloaded at CBS website.
A tcsh interpreter is used to run ProP and needs to be installed.
apt–get install tcsh
The script prop needs to be edited to set the full path to prop folder.
setenv PROPHOME /var/www/genotate/services/prop
TMHMM predicts of transmembrane domains, which fundamentally rule all the membrane biochemical processes, with the hidden Markov models. TMHMM can be downloaded at CBS website.
The script tmhmm need to be edited to set the full path to tmhmm folder.
$opt_basedir = "/var/www/genotate/services/tmhmm/";
Check the path to perl /usr/bin/perl or /usr/local/bin/perl in the files tmhmm/bin/tmhmm and tmhmm/bin/tmhmmformat.pl. If no results are printed by tmhmm, the option –d enable tmhmm error messages.
SIGNALP predicts the secretory signal peptide, a ubiquitous signal that targets for translocation across the membrane, based on neural network. SignalP can be downloaded at CBS website.
The script signalp need to be edited to set the full path to signalP folder.
$ENV{SIGNALP} = '/var/www/genotate/services/signalp';
Glycosylation covalently attach a carbohydrate to proteins and lipids. Some proteins require to be glycosylated to fold correctly. NetCGlyc produces neural network predictions of C–mannosylation sites in mammalian proteins. NETCGLYC and is available on CBS website. The script which launch NETCGLYC need to be edited to set the full path to the installation directory.
NetNglyc predicts N–Glycosylation sites in human proteins using artificial neural networks. NETNGLYC and is available on CBS website. The script which launch NETNGLYC need to be edited to set the full path to the installation directory.
Predict the location of linear B–cell epitopes using a combination of a hidden Markov model and a propensity scale method. MHC–II and is available on CBS website. The script which launch BEPIPRED need to be edited to set the full path to the installation directory.
Determine each subsequence's ability to bind to a specific MHC class I molecule. MHC–I and is available on IEDB website. A configuration script is available in the tool directory.
wget http://media.iedb.org/tools/mhci/latest/IEDB_MHC_I–2.15.4.tar.gz
cd mhc_i
./configure
#if error ImportError: No module named pkg_resources
#sudo apt–get install ––reinstall python–pkg–resources
After installation, remove the line waiting for input in predict_binding.py as following.
#if not sys.stdin.isatty():
# stdin = sys.stdin.readline().strip()
# args.append(stdin)
Predict MHC Class II epitopes, including a consensus approach which combines NN–align, SMM–align and Combinatorial library methods. MHC–II and is available on IEDB website. A configuration script is available in the tool directory.
apt–get install gawk
wget http://media.iedb.org/tools/mhcii/latest/IEDB_MHC_II–2.16.2.tar.gz
cd mhc_ii
./configure.py
#if error ImportError: No module named pkg_resources
#sudo apt–get install ––reinstall python–pkg–resources
After installation, remove the line waiting for input in predict_binding.py as following.
#if not sys.stdin.isatty():
# stdin = sys.stdin.readline().strip()
# sys.argv.append(stdin)
wget http://eddylab.org/software/hmmer/2.3/hmmer-2.3.tar.gz
cd hmmer-2.3
cpan install Getopt::Long
./configure
make
make check
make install
mkdir rnammer
wget [rnammer link from CBS]
tar -xzvf rnammer.tar.Z
libxml-parser-perl
cpan install XML::Simple
edit rnammer paths
edit core-rnammer remove --cpu 1
wget http://lowelab.ucsc.edu/software/tRNAscan-SE.tar.gz
tar -zxvf tRNAscan-SE.tar.gz
mv tRNAscan-SE-1.3.1 tRNAscan-SE
cd tRNAscan-SE
edit makefile paths
make
make install
Please verify that the services can be executed.
The following files and folders shoul be executable:
- Java executables in jdk/bin and jdk/jre/bin
- BLAST executables in blast/bin
- ProP executables in prop/bin and prop/how and the script prop
- TMHMM executables in tmhmm/bin and the script tmhmm
- SignalP executables in signalP/bin and the script signalp
Genotate java executable is available on Github. Genotate can be downloaded using the following command:
wget https://github.com/tchitchek–lab/genotate/blob/master/binaries/genotate.jar
Genotate requires a configuration file with the path to all programs and databases used for the annotation. Please ensure the file genotate.config is in genotate.jar folder.
Below you can found an example for the configuration file, available in github.
BEPIPRED:/var/www/genotate.life/services/bepipred/bepipred
BLAST:/var/www/genotate.life/services/blast/bin
BLASTALL:/var/www/genotate.life/services/interproscan-5.30-69.0/bin/blast/2.2.24
BLASTDB:/var/www/genotate.life/workspace/blastdb
INTERPROSCAN:/var/www/genotate.life/services/interproscan-5.30-69.0/interproscan.sh
JAVA:/var/www/genotate.life/services/javajdk/bin
MHCI:/var/www/genotate.life/services/mhc_i/src/predict_binding.py
MHCII:/var/www/genotate.life/services/mhc_ii/mhc_II_binding.py
NETCGLYC:/var/www/genotate.life/services/netCglyc/netCglyc
NETNGLYC:/var/www/genotate.life/services/netNglyc/netNglyc
PROP:/var/www/genotate.life/services/prop/prop
SIGNALP:/var/www/genotate.life/services/signalp/signalp
TMHMM:/var/www/genotate.life/services/tmhmm/bin/tmhmm
RNAMMER:/var/www/genotate.life/services/rnammer/rnammer
TRNASCANSE:/var/www/genotate.life/services/trnascanse/tRNAscan-SE
TRNASCANSE_ENV:/var/www/genotate.life/services/trnascanse/setup.tRNAscan-SE
Genotate can be executed with the following commands:
java –jar genotate.jar –input example.fasta –output test –services TMHMM
Multiple options are available to run Genotate.
-input Input nucleic fasta file path
-output Output folder path
-services Services to run
-services_messages Display the services messages
-inner_orf Allows orf contained in larger ones
-outside_orf Allows partial orf lacking either a codon stop or a codon start
-orf_min_size 100 Filter orf to keep only those long enough. The size is in nucleic bases
-region_by_run 100 Number of region computed together
-refresh_time 10 Waiting time in seconds between each results check
-threads 8 Number of jobs computed at the same time
-start_codon ATG Start codon(s) used to search for orf
-stop_codon TAG,TGA,TAA Stop codon(s) used to search for orf
-ignore_reverse Do not compute the annotation on the reverse strand
-ignore_ncrna Do not compute the annotation of ncrna
-checkORF Use CPAT to check the ORF coding potential
-checkORF_threshold 0.5 Set CPAT threshold
Multiple services and databases are available to run Genotate functinal annotation. For each service a score is available to control the quality of the annotations.
BLASTN [database,identity,query cover,subject cover] by default 85,50,50 (min 0 to max 100)
BLASTP [database,identity,query cover,subject cover] by default 85,50,50 (min 0 to max 100)
BEPIPRED [score] by default threshold = 0.5 (min 0 to max 1)
MHCI [score] by default threshold = 1 (min 0 to max 2)
MHCII [score] by default threshold = 1 (min 0 to max 2)
NETCGLYC [score] by default threshold = 0.5 (min 0 to max 1)
NETNGLYC [score] by default threshold = 0.5 (min 0 to max 1)
PROP [score] by default threshold = 0.2 (min 0 to max 1)
SIGNALP [score] by default threshold = 0.45(min 0 to max 1)
CDD [evalue] by default evalue = 0.05 (min 0 to max 1)
COILS [evalue] by default evalue = 0.05 (min 0 to max 1)
GENE3D [evalue] by default evalue = 0.05 (min 0 to max 1)
HAMAP [evalue] by default evalue = 0.05 (min 0 to max 1)
MOBIDBLITE [evalue] by default evalue = 0.05 (min 0 to max 1)
PANTHER [evalue] by default evalue = 0.05 (min 0 to max 1)
PFAM [evalue] by default evalue = 0.05 (min 0 to max 1)
PIRSF [evalue] by default evalue = 0.05 (min 0 to max 1)
PRINTS [evalue] by default evalue = 0.05 (min 0 to max 1)
PRODOM [evalue] by default evalue = 0.05 (min 0 to max 1)
PROSITEPATTERNS [evalue] by default evalue = 0.05 (min 0 to max 1)
PROSITEPROFILES [evalue] by default evalue = 0.05 (min 0 to max 1)
SFLD [evalue] by default evalue = 0.05 (min 0 to max 1)
SMART [evalue] by default evalue = 0.05 (min 0 to max 1)
SUPERFAMILY [evalue] by default evalue = 0.05 (min 0 to max 1)
TIGRFAM [evalue] by default evalue = 0.05 (min 0 to max 1)
TMHMM no scores availables
Genotate provides the annotation results in multiple results files.
transcript.fasta: this file contains all the transcripts used in input.
>ID1 Description
ATGGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGTAA
>ID2 Description
ATGGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGTAA
>ID3 Description
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
transcript_clean.fasta: this file contains conserved for annotations.
>ID1 Description
ATGGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGTAA
>ID2 Description
ATGGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGTAA
transcript_info.tab: The transcript informations.
transcript_id transcript_name transcript_desc transcript_size
0 AL049998.1 Homo sapiens mRNA; cDNA DKFZp564L222 Phosphatidylinositol–4–phosphate 3–kinase 1304
1 NM_002645.3 Homo sapiens phosphatidylinositol–4–phosphate 3–kinase, mRNA 2936
region_nucl.fasta and region_prot.fasta contains the coding and noncoding nucleic sequences and for the coding regions the translated protein sequences.
>Region1 Description
ATGGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGTAA
>Region2 Description
ATGGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCC
CCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGCCCGGGTAA
region_info.tab: The region informations, with the positions on the transcript and the strand, the size and the type which can be set to inner or outside.
region_id begin end size strand coding type transcript_id
0 423 609 186 + coding 0
1 669 888 219 + coding 0
2 294 423 129 + coding 0
3 717 834 117 + coding 0
4 249 351 102 + coding 1
5 213 2919 2706 + coding 1
6 666 774 108 + coding 1
7 2169 2340 171 + coding 1
8 966 1230 264 – coding 1
all_annotations.tab: The annotations with the region_id, the service used, the position of the annotation on the region, the name and the description of the annotation.
region_id service begin end name description
10 TMHMM 3 114 outside
10 PROP 27 54 cleavage site: VSGSVKRGV
9 TMHMM 3 120 inside
8 TMHMM 3 261 inside
7 TMHMM 3 168 outside
6 TMHMM 3 105 outside
5 TMHMM 3 2703 outside
5 CDD 2031 2547 cd04012 C2A_PI3K_class_II
4 TMHMM 3 99 inside
4 PROP 33 60 cleavage site: KRCGQRRSI
For each region an SVG graph is generated (which can be opened in a web browser) and allow to visualize the annotations on the ORF.


