###### Protein

lamindb provides access to the following public Protein ontologies
through bionty:

1. Uniprot

Here we show how to access and search Protein ontologies to
standardize new data.

 import bionty as bt
 import pandas as pd

##### PublicOntology objects

Let us create a public ontology accessor with ".public" method, which
chooses a default public ontology source from "Source". It's a
PublicOntology object, which you can think about as a public registry:

 proteins = bt.Protein.public(organism="human")
 proteins

As for registries, you can export the ontology as a "DataFrame":

 df = proteins.to_dataframe()
 df.head()

Unlike registries, you can also export it as a Pronto object via
"public.ontology".

##### Look up terms

As for registries, terms can be looked up with auto-complete:

 lookup = proteins.lookup()

The "." accessor provides normalized terms (lower case, only contains
alphanumeric characters and underscores):

 lookup.ac3

To look up the exact original strings, convert the lookup object to
dict and use the "[]" accessor:

 lookup_dict = lookup.dict()
 lookup_dict["AC3"]

By default, the "name" field is used to generate lookup keys. You can
specify another field to look up:

 lookup = proteins.lookup(proteins.gene_symbol)

 lookup.rab4a

##### Search terms

Search behaves in the same way as it does for registries:

 proteins.search("RAS").head(3)

By default, search also covers synonyms and all other fields
containing strings:

 proteins.search("member of RAS oncogene family like 2B").head(3)

Search specific field (by default, search is done on all fields
containing strings):

 proteins.search(
 "RABL2B",
 field=proteins.gene_symbol,
 ).head()

##### Standardize Protein identifiers

Let us generate a "DataFrame" that stores a number of Protein
identifiers, some of which corrupted:

 df_orig = pd.DataFrame(
 index=[
 "A0A024QZ08",
 "X6RLV5",
 "X6RM24",
 "A0A024QZQ1",
 "This protein does not exist",
 ]
 )
 df_orig

We can check whether any of our values are validated against the
ontology reference:

 validated = proteins.validate(df_orig.index, proteins.name)
 df_orig.index[~validated]

##### Ontology source versions

For any given entity, we can choose from a number of versions:

 bt.Source.filter(entity="bionty.Protein").to_dataframe()

 # only lists the sources that are currently used
 bt.Source.filter(entity="bionty.Protein", currently_used=True).to_dataframe()

When instantiating a Bionty object, we can choose a source or version:

 source = bt.Source.filter(
 name="uniprot", organism="human"
 ).first()
 proteins= bt.Protein.public(source=source)
 proteins

The currently used ontologies can be displayed using:

 bt.Source.filter(currently_used=True).to_dataframe()