Data Methods
This page describes the data collection and processing methods used in the Ortholog Search and Experimental modes of the Ortholog Finder Tool. The methodology and datasets reflect the state of the databases at the time of the tool's original publication in 2016. Ortholog mappings, pathway annotations, and protein interaction data were retrieved from the sources listed below during 2013-2015.
S. cerevisiae / Budding yeast
We used 451
S. cerevisiae gene strains, from .
We checked then we found that there are , and gene strains.
Therefore we corrected them and used a gene list, which contained 445
gene strains.
S. cerevisiae gene strains, from .We checked then we found that there are , and gene strains.
Therefore we corrected them and used a gene list, which contained 445
gene strains.
We expanded this gene list with the research performed by .
This research was a wide pharmaco-epistasis analysis, which mainly based on the research above.
Finally we added protein to our gene list.
Thus we created the final S. cerevisiae gene strain list, which contained 489
probable size control genes.
Till this point we used as primary key the Yeast Genome Database's systematic ORF IDs
This research was a wide pharmaco-epistasis analysis, which mainly based on the research above.
Finally we added protein to our gene list.
Thus we created the final S. cerevisiae gene strain list, which contained 489
probable size control genes. Till this point we used as primary key the Yeast Genome Database's systematic ORF IDs
We used (version 3.2.106, December 2013) to retrieve this gene list's protein-protein network.
Therefore we needed gene/protein mapping to use this database for research, because BioGRID used NCBI Entrez Gene IDs as primary keys.
We mapped all
but one
YGD ORF IDs to UniProtKB IDs. Later UniProtKB IDs were mapped to Entrez Gene IDs.
Our query mapped 450 UniProtKB IDs to 476
(462 unique) Entrez IDs, while 40
UniProtKB IDs were not found to have an Entrez ID pair.
During protein mapping we used the UniProt.org's ID Mapper API.
Therefore we needed gene/protein mapping to use this database for research, because BioGRID used NCBI Entrez Gene IDs as primary keys.
We mapped all
but one
YGD ORF IDs to UniProtKB IDs. Later UniProtKB IDs were mapped to Entrez Gene IDs. Our query mapped 450 UniProtKB IDs to 476
(462 unique) Entrez IDs, while 40
UniProtKB IDs were not found to have an Entrez ID pair.During protein mapping we used the UniProt.org's ID Mapper API.
S. pombe / Fission yeast
To find orthologs we used PomBase manually curated one S. pombe to one S. cerevisiae ortholog list and .
First we investigated the PomBase ortholog lists.
We found that from the 489 S. cerevisiae genes (YGD ORF IDs) 392 orthologs
had an ortholog, while 97 had no orthologs.
These genes had 507 S. cerevisiae - S. pombe ortholog pairs, from which 477
S. pombe were unique.
In this point we used PomBase Systematic ORF id as primary key for S. pombe orthologs.
We found that from the 489 S. cerevisiae genes (YGD ORF IDs) 392 orthologs
had an ortholog, while 97 had no orthologs. These genes had 507 S. cerevisiae - S. pombe ortholog pairs, from which 477
S. pombe were unique. In this point we used PomBase Systematic ORF id as primary key for S. pombe orthologs.
Later we mapped these IDs to UniProtKB IDs to be comparable with inParanoid database, we used PomBase's gp2swiss mappings.
We mapped the 477
PomBase ORF IDs to 495 UniProt IDs, while 476
IDs were unique and 7 were not found.
We mapped the 477
PomBase ORF IDs to 495 UniProt IDs, while 476
IDs were unique and 7 were not found.
To retrieve connections for S. pombe we mapped UniProt IDs to Entrez Gene IDs with UniProt.org's ID Mapper API.
We mapped all of the 476 UniProt IDs to 539 Entrez IDs, while 470
IDs were unique.
We mapped all of the 476 UniProt IDs to 539 Entrez IDs, while 470
IDs were unique.
Second we made a query from inParanoid.
We found from the 490 S. cerevisiae proteins (UniProtKB IDs) 332 proteins has an ortholog (158 proteins
were not found).
In S. pombe 388 (all percent confidence) orthologs
were found (356 unique UniProt IDs).
To retrieve connections for S. pombe we mapped all of the 356 UniProt IDs to 370 Entrez IDs
(368 unique IDs).
We found from the 490 S. cerevisiae proteins (UniProtKB IDs) 332 proteins has an ortholog (158 proteins
were not found). In S. pombe 388 (all percent confidence) orthologs
were found (356 unique UniProt IDs). To retrieve connections for S. pombe we mapped all of the 356 UniProt IDs to 370 Entrez IDs
(368 unique IDs).
H. sapiens / Human
To find human orthologs we could also use PomBase ortholog list and inParanoid database.
We started the investigation with PomBase ortholog list. This is a manually curated list between S. pombe and H. sapiens.
We used for the primary query list the S. pombe's PomBase orthologs of S. cerevisiae's genes.
We found from the 477 S. pombe genes (PomBase ORF IDs) 397 orthologs
(regular name).
These proteins had 606 human orthologs, from which 469 were unique.
We used for the primary query list the S. pombe's PomBase orthologs of S. cerevisiae's genes.
We found from the 477 S. pombe genes (PomBase ORF IDs) 397 orthologs
(regular name). These proteins had 606 human orthologs, from which 469 were unique.
Later we made a query from inParanoid.
We found from the 490 S. cerevisiae proteins (UniProtKB IDs) 243 proteins has an orthologs (247 proteins
were not found).
In H. sapiens 351 (all percent confidence) orthologs
were found (328 unique UniProt IDs).
To retrieve connections for H. sapiens we mapped 321 of the 328 UniProt IDs to 328 Entrez IDs
(326 unique IDs, 7 were not found
).
We found from the 490 S. cerevisiae proteins (UniProtKB IDs) 243 proteins has an orthologs (247 proteins
were not found). In H. sapiens 351 (all percent confidence) orthologs
were found (328 unique UniProt IDs). To retrieve connections for H. sapiens we mapped 321 of the 328 UniProt IDs to 328 Entrez IDs
(326 unique IDs, 7 were not found
).
A. thaliana / Thale cress
For Arabidopsis we did not have a manually curated ortholog source so we retrieved our query only from inParanoid.
We found from the 490 S. cerevisiae proteins (UniProtKB IDs) 239 proteins has an ortholog (251 proteins
were not found).
In A. thaliana 783 (all percent confidence) orthologs
were found (738 unique UniProt IDs).
To retrieve connections for A. thaliana we mapped all of the 738 UniProt IDs to 759 Entrez IDs
(755 unique IDs).
were not found). In A. thaliana 783 (all percent confidence) orthologs
were found (738 unique UniProt IDs). To retrieve connections for A. thaliana we mapped all of the 738 UniProt IDs to 759 Entrez IDs
(755 unique IDs).