Phenotype

lamindb provides access to the following public protein ontologies through bionty:

  1. Human Phenotype

  2. Phecodes

  3. PATO

  4. Mammalian Phenotype

Here we show how to access and search Phenotype ontologies to standardize new data.

import bionty as bt
import pandas as pd
💡 connected lamindb: testuser1/test-public-ontologies

PublicOntology objects

Let us create a public ontology accessor with .public method, which chooses a default public ontology source from PublicSource. It’s a PublicOntology object, which you can think about as a public registry:

phenotypes = bt.Phenotype.public(organism="human")
phenotypes
PublicOntology
Entity: Phenotype
Organism: human
Source: hp, 2024-03-06
#terms: 18697

As for registries, you can export the ontology as a DataFrame:

df = phenotypes.df()
df.head()
name definition synonyms parents
ontology_id
HP:0000001 All None None []
HP:0000002 Abnormality of body height Deviation From The Norm Of Height With Respect... Abnormality of body height [HP:0001507]
HP:0000003 Multicystic kidney dysplasia Multicystic Dysplasia Of The Kidney Is Charact... Multicystic kidneys|Multicystic dysplastic kid... [HP:0000107]
HP:0000005 Mode of inheritance The Pattern In Which A Particular Genetic Trai... Inheritance [HP:0000001]
HP:0000006 Autosomal dominant inheritance A Mode Of Inheritance That Is Observed For Tra... monoallelic_autosomal|Autosomal dominant [HP:0034345]

Unlike registries, you can also export it as a Pronto object via public.ontology.

Look up terms

As for registries, terms can be looked up with auto-complete:

lookup = phenotypes.lookup()

The . accessor provides normalized terms (lower case, only contains alphanumeric characters and underscores):

lookup.eeg_with_persistent_abnormal_rhythmic_activity
Phenotype(ontology_id='HP:0010846', name='EEG with persistent abnormal rhythmic activity', definition=None, synonyms='EEG: persistent abnormal rhythmic activity', parents=array(['HP:0011176'], dtype=object))

To look up the exact original strings, convert the lookup object to dict and use the [] accessor:

lookup_dict = lookup.dict()
lookup_dict["EEG with persistent abnormal rhythmic activity"]
Phenotype(ontology_id='HP:0010846', name='EEG with persistent abnormal rhythmic activity', definition=None, synonyms='EEG: persistent abnormal rhythmic activity', parents=array(['HP:0011176'], dtype=object))

By default, the name field is used to generate lookup keys. You can specify another field to look up:

lookup = phenotypes.lookup(phenotypes.ontology_id)
lookup.hp_0000003
Phenotype(ontology_id='HP:0000003', name='Multicystic kidney dysplasia', definition='Multicystic Dysplasia Of The Kidney Is Characterized By Multiple Cysts Of Varying Size In The Kidney And The Absence Of A Normal Pelvicaliceal System. The Condition Is Associated With Ureteral Or Ureteropelvic Atresia, And The Affected Kidney Is Nonfunctional.', synonyms='Multicystic kidneys|Multicystic dysplastic kidney|Multicystic renal dysplasia', parents=array(['HP:0000107'], dtype=object))

Search terms

Search behaves in the same way as it does for registries:

phenotypes.search("dysplasia").head(3)
ontology_id definition synonyms parents __ratio__
name
Hip dysplasia HP:0001385 The Presence Of Developmental Dysplasia Of The... DDH|Developmental dysplasia of the hip|Congeni... [HP:0003272] 95.0
Focal cortical dysplasia type III HP:0032054 A Type Of Focal Cortical Dysplasia That Is Cha... None [HP:0032046] 90.0
Polycystic kidney dysplasia HP:0000113 The Presence Of Multiple Cysts In Both Kidneys. Enlarged polycystic kidneys|Polycystic kidneys [HP:0000107] 90.0

By default, search also covers synonyms:

phenotypes.search("Congenital hip dysplasia").head(3)
ontology_id definition synonyms parents __ratio__
name
Hip dysplasia HP:0001385 The Presence Of Developmental Dysplasia Of The... DDH|Developmental dysplasia of the hip|Congeni... [HP:0003272] 100.000000
Congenital hip dislocation HP:0001374 None Congenital hip anomaly|Congenital dislocation ... [HP:0002827, HP:0001385] 80.000000
Congenital alveolar dysplasia HP:0033210 Arrest Of Lung Development In The Cananicular ... None [HP:0006703] 79.245283

You can turn this off synonym by passing synonyms_field=None:

phenotypes.search("Congenital hip dysplasia", synonyms_field=None).head(3)
ontology_id definition synonyms parents __ratio__
name
Congenital hip dislocation HP:0001374 None Congenital hip anomaly|Congenital dislocation ... [HP:0002827, HP:0001385] 80.000000
Congenital alveolar dysplasia HP:0033210 Arrest Of Lung Development In The Cananicular ... None [HP:0006703] 79.245283
Toenail dysplasia HP:0100797 An Abnormality Of The Development Of The Toena... Abnormal toenail development|Dysplastic toenails [HP:0008388, HP:0002164] 73.170732

Search another field (default is .name):

phenotypes.search(
    "lack of development of speech and language",
    field=phenotypes.definition,
).head()
ontology_id name synonyms parents __ratio__
definition
Complete Lack Of Development Of Speech And Language Abilities. HP:0001344 Absent speech Lack of language development|Nonverbal|No spee... [HP:0002167, HP:0000750] 81.553398
Lack Of Development Of One Lung. HP:0030707 Unilateral lung agenesis Unilateral pulmonary agenesis [HP:0006703] 76.712329
Absence Of The Nasal Bone. HP:0010941 Aplasia of the nasal bone Failure of development of the nasal bone|Lack ... [HP:0010940] 73.417722
Agenesis Of Canine Tooth. HP:0012738 Agenesis of canine Failure of development of canine|Absent canine... [HP:0001592, HP:0011078] 67.567568
Agenesis Of Secondary Molar Tooth. HP:0011055 Agenesis of permanent molar Agenesis of secondary molar|Failure of develop... [HP:0011054] 67.469880

Standardize Phenotype identifiers

Let us generate a DataFrame that stores a number of Phenotype identifiers, some of which corrupted:

df_orig = pd.DataFrame(
    index=[
        "Specific learning disability",
        "Dystonia",
        "Cerebral hemorrhage",
        "Slurred speech",
        "This phenotype does not exist",
    ]
)
df_orig
Specific learning disability
Dystonia
Cerebral hemorrhage
Slurred speech
This phenotype does not exist

We can check whether any of our values are validated against the ontology reference:

validated = phenotypes.validate(df_orig.index, phenotypes.name)
df_orig.index[~validated]
4 terms (80.00%) are validated
1 term (20.00%) is not validated: This phenotype does not exist
Index(['This phenotype does not exist'], dtype='object')

Ontology source versions

For any given entity, we can choose from a number of versions:

bt.PublicSource.filter(entity="Phenotype").df()
uid entity organism currently_used source source_name version url md5 source_website run_id created_by_id updated_at
id
54 2WLc Phenotype human True hp Human Phenotype Ontology 2024-03-06 https://github.com/obophenotype/human-phenotyp... 36b0d00c24a68edb9131707bc146a4c7 https://hpo.jax.org None 1 2024-06-19 23:14:43.639132+00:00
55 6jHz Phenotype human False hp Human Phenotype Ontology 2023-06-17 https://github.com/obophenotype/human-phenotyp... 65e8d96bc81deb893163927063b10c06 https://hpo.jax.org None 1 2024-06-19 23:14:43.639228+00:00
56 5Tzl Phenotype human False hp Human Phenotype Ontology 2023-04-05 https://github.com/obophenotype/human-phenotyp... bdf866e11d37cf6fd2aef25c325b2c8a https://hpo.jax.org None 1 2024-06-19 23:14:43.639324+00:00
57 5EQM Phenotype human False hp Human Phenotype Ontology 2023-01-27 https://github.com/obophenotype/human-phenotyp... ceeb3ada771908deef620d74cd8e6b0f https://hpo.jax.org None 1 2024-06-19 23:14:43.639419+00:00
58 6zE1 Phenotype mammalian True mp Mammalian Phenotype Ontology 2024-02-07 https://github.com/mgijax/mammalian-phenotype-... 31c27ed2c7d5774f8b20a77e4e1fd278 https://github.com/mgijax/mammalian-phenotype-... None 1 2024-06-19 23:14:43.639514+00:00
59 4q5A Phenotype mammalian False mp Mammalian Phenotype Ontology 2023-05-31 https://github.com/mgijax/mammalian-phenotype-... be89052cf6d9c0b6197038fe347ef293 https://github.com/mgijax/mammalian-phenotype-... None 1 2024-06-19 23:14:43.639610+00:00
60 7EnA Phenotype zebrafish True zp Zebrafish Phenotype Ontology 2024-01-22 https://github.com/obophenotype/zebrafish-phen... 01600a5d392419b27fc567362d4cfff8 https://github.com/obophenotype/zebrafish-phen... None 1 2024-06-19 23:14:43.639706+00:00
61 6Czy Phenotype zebrafish False zp Zebrafish Phenotype Ontology 2022-12-17 https://github.com/obophenotype/zebrafish-phen... 03430b567bf153216c0fa4c3440b3b24 https://github.com/obophenotype/zebrafish-phen... None 1 2024-06-19 23:14:43.639803+00:00
62 5Qpm Phenotype human False phe Phecodes ICD10 map 1.2 s3://bionty-assets/df_human__phe__1.2__Phenoty... 741033ee1b13df7c41b4849e8bd02f13 https://phewascatalog.org/phecodes_icd10 None 1 2024-06-19 23:14:43.639900+00:00
63 55lY Phenotype all True pato Phenotype And Trait Ontology 2023-05-18 http://purl.obolibrary.org/obo/pato/releases/2... bd472f4971492109493d4ad8a779a8dd https://github.com/pato-ontology/pato None 1 2024-06-19 23:14:43.639997+00:00

When instantiating a Bionty object, we can choose a source or version:

public_source = bt.PublicSource.filter(
    source="hp", version="2023-06-17", organism="human"
).one()
phenotypes= bt.Phenotype.public(public_source=public_source)
phenotypes
❗ loading non-default source inside a LaminDB instance
PublicOntology
Entity: Phenotype
Organism: human
Source: hp, 2023-06-17
#terms: 17653

The currently used ontologies can be displayed using:

bt.PublicSource.filter(currently_used=True).df()
Hide code cell output
uid entity organism currently_used source source_name version url md5 source_website run_id created_by_id updated_at
id
1 5Dlc Organism vertebrates True ensembl Ensembl release-112 https://ftp.ensembl.org/pub/release-112/specie... 0ec37e77f4bc2d0b0b47c6c62b9f122d https://www.ensembl.org None 1 2024-06-19 23:14:43.633921+00:00
6 2Jzh Organism bacteria True ensembl Ensembl release-57 https://ftp.ensemblgenomes.ebi.ac.uk/pub/bacte... ee28510ed5586ea7ab4495717c96efc8 https://www.ensembl.org None 1 2024-06-19 23:14:43.634460+00:00
7 1kdI Organism fungi True ensembl Ensembl release-57 http://ftp.ensemblgenomes.org/pub/fungi/releas... dbcde58f4396ab8b2480f7fe9f83df8a https://www.ensembl.org None 1 2024-06-19 23:14:43.634557+00:00
8 2mIM Organism metazoa True ensembl Ensembl release-57 http://ftp.ensemblgenomes.org/pub/metazoa/rele... 424636a574fec078a61cbdddb05f9132 https://www.ensembl.org None 1 2024-06-19 23:14:43.634654+00:00
9 2XQ6 Organism plants True ensembl Ensembl release-57 https://ftp.ensemblgenomes.ebi.ac.uk/pub/plant... eadaa1f3e527e4c3940c90c7fa5c8bf4 https://www.ensembl.org None 1 2024-06-19 23:14:43.634750+00:00
10 1Vzs Organism all True ncbitaxon NCBItaxon Ontology 2023-06-20 s3://bionty-assets/df_all__ncbitaxon__2023-06-... 00d97ba65627f1cd65636d2df22ea76c https://github.com/obophenotype/ncbitaxon None 1 2024-06-19 23:14:43.634846+00:00
11 1hx4 Gene human True ensembl Ensembl release-112 s3://bionty-assets/df_human__ensembl__release-... 4ccda4d88720a326737376c534e8446b https://www.ensembl.org None 1 2024-06-19 23:14:43.634944+00:00
15 76FX Gene mouse True ensembl Ensembl release-112 s3://bionty-assets/df_mouse__ensembl__release-... 519cf7b8acc3c948274f66f3155a3210 https://www.ensembl.org None 1 2024-06-19 23:14:43.635343+00:00
19 7LW6 Gene saccharomyces cerevisiae True ensembl Ensembl release-112 s3://bionty-assets/df_saccharomyces cerevisiae... 11775126b101233525a0a9e2dd64edae https://www.ensembl.org None 1 2024-06-19 23:14:43.635731+00:00
22 7llW Protein human True uniprot Uniprot 2023-03 s3://bionty-assets/df_human__uniprot__2023-03_... 1c46e85c6faf5eff3de5b4e1e4edc4d3 https://www.uniprot.org None 1 2024-06-19 23:14:43.636031+00:00
24 5U7J Protein mouse True uniprot Uniprot 2023-03 s3://bionty-assets/df_mouse__uniprot__2023-03_... 9d5e9a8225011d3218e10f9bbb96a46c https://www.uniprot.org None 1 2024-06-19 23:14:43.636233+00:00
26 5nkB CellMarker human True cellmarker CellMarker 2.0 s3://bionty-assets/human_cellmarker_2.0_CellMa... d565d4a542a5c7e7a06255975358e4f4 http://bio-bigdata.hrbmu.edu.cn/CellMarker None 1 2024-06-19 23:14:43.636426+00:00
27 6AFz CellMarker mouse True cellmarker CellMarker 2.0 s3://bionty-assets/mouse_cellmarker_2.0_CellMa... 189586732c63be949e40dfa6a3636105 http://bio-bigdata.hrbmu.edu.cn/CellMarker None 1 2024-06-19 23:14:43.636524+00:00
28 6cbC CellLine all True clo Cell Line Ontology 2022-03-21 https://data.bioontology.org/ontologies/CLO/su... ea58a1010b7e745702a8397a526b3a33 https://bioportal.bioontology.org/ontologies/CLO None 1 2024-06-19 23:14:43.636619+00:00
29 3DeN CellType all True cl Cell Ontology 2024-02-13 http://purl.obolibrary.org/obo/cl/releases/202... https://obophenotype.github.io/cell-ontology None 1 2024-06-19 23:14:43.636715+00:00
34 1AyH Tissue all True uberon Uberon multi-species anatomy ontology 2024-02-20 http://purl.obolibrary.org/obo/uberon/releases... 2048667b5fdf93192384bdf53cafba18 http://obophenotype.github.io/uberon None 1 2024-06-19 23:14:43.637192+00:00
39 LoCG Disease all True mondo Mondo Disease Ontology 2024-02-06 http://purl.obolibrary.org/obo/mondo/releases/... 78914fa236773c5ea6605f7570df6245 https://mondo.monarchinitiative.org None 1 2024-06-19 23:14:43.637670+00:00
44 2mou Disease human True doid Human Disease Ontology 2024-01-31 http://purl.obolibrary.org/obo/doid/releases/2... b36c15a4610757094f8db64b78ae2693 https://disease-ontology.org None 1 2024-06-19 23:14:43.638168+00:00
51 4usY ExperimentalFactor all True efo The Experimental Factor Ontology 3.63.0 http://www.ebi.ac.uk/efo/releases/v3.63.0/efo.owl 603e6f6981d53d501c5921aa3940b095 https://bioportal.bioontology.org/ontologies/EFO None 1 2024-06-19 23:14:43.638843+00:00
54 2WLc Phenotype human True hp Human Phenotype Ontology 2024-03-06 https://github.com/obophenotype/human-phenotyp... 36b0d00c24a68edb9131707bc146a4c7 https://hpo.jax.org None 1 2024-06-19 23:14:43.639132+00:00
58 6zE1 Phenotype mammalian True mp Mammalian Phenotype Ontology 2024-02-07 https://github.com/mgijax/mammalian-phenotype-... 31c27ed2c7d5774f8b20a77e4e1fd278 https://github.com/mgijax/mammalian-phenotype-... None 1 2024-06-19 23:14:43.639514+00:00
60 7EnA Phenotype zebrafish True zp Zebrafish Phenotype Ontology 2024-01-22 https://github.com/obophenotype/zebrafish-phen... 01600a5d392419b27fc567362d4cfff8 https://github.com/obophenotype/zebrafish-phen... None 1 2024-06-19 23:14:43.639706+00:00
63 55lY Phenotype all True pato Phenotype And Trait Ontology 2023-05-18 http://purl.obolibrary.org/obo/pato/releases/2... bd472f4971492109493d4ad8a779a8dd https://github.com/pato-ontology/pato None 1 2024-06-19 23:14:43.639997+00:00
64 48aa Pathway all True go Gene Ontology 2023-05-10 https://data.bioontology.org/ontologies/GO/sub... e9845499eadaef2418f464cd7e9ac92e http://geneontology.org None 1 2024-06-19 23:14:43.640094+00:00
67 3rm9 BFXPipeline all True lamin Bioinformatics Pipeline 1.0.0 s3://bionty-assets/bfxpipelines.json a7eff57a256994692fba46e0199ffc94 https://lamin.ai None 1 2024-06-19 23:14:43.640389+00:00
68 5alK Drug all True dron Drug Ontology 2024-03-02 https://data.bioontology.org/ontologies/DRON/s... 84138459de4f65034e979f4e46783747 https://bioportal.bioontology.org/ontologies/DRON None 1 2024-06-19 23:14:43.640486+00:00
70 7CRn DevelopmentalStage human True hsapdv Human Developmental Stages 2020-03-10 http://aber-owl.net/media/ontologies/HSAPDV/11... 52181d59df84578ed69214a5cb614036 https://github.com/obophenotype/developmental-... None 1 2024-06-19 23:14:43.640679+00:00
71 16tR DevelopmentalStage mouse True mmusdv Mouse Developmental Stages 2020-03-10 http://aber-owl.net/media/ontologies/MMUSDV/9/... 5bef72395d853c7f65450e6c2a1fc653 https://github.com/obophenotype/developmental-... None 1 2024-06-19 23:14:43.640781+00:00
72 3Tlc Ethnicity human True hancestro Human Ancestry Ontology 3.0 https://github.com/EBISPOT/hancestro/raw/3.0/h... 76dd9efda9c2abd4bc32fc57c0b755dd https://github.com/EBISPOT/hancestro None 1 2024-06-19 23:14:43.642901+00:00
73 5JnV BioSample all True ncbi NCBI BioSample attributes 2023-09 s3://bionty-assets/df_all__ncbi__2023-09__BioS... 918db9bd1734b97c596c67d9654a4126 https://www.ncbi.nlm.nih.gov/biosample/docs/at... None 1 2024-06-19 23:14:43.643010+00:00