Online program

Arabidopsis GeneCloud — semantic terms over-represented in the description of a gene list

Paste a list of Arabidopsis genes. GeneCloud reads what is written about every gene — TAIR curator summaries and descriptions, mutant phenotypes, UniProt function, NCBI Gene and Gene Ontology — and returns the words and phrases that are significantly more frequent in your list than in the background, as a word cloud in which the size of a word is its statistical enrichment. It complements GO analysis: it finds vocabulary that no ontology term captures. The analysis runs in your browser: your gene list is never sent anywhere.

0 identifiers
Examples from the GeneCloud paper: · ·
Background *Required: every gene that could have been in your list
Annotation text
Concepts
Keep terms occurring in at least
FDR threshold
p.adjust method
Cloudcan be changed after the analysis
word size:

Results

fold enrichment
wordtypecountbackground enrich. ratiop-valueFDRmerged with

Click a term to list the genes that carry it, with their annotation (the term is highlighted). Click a column to sort.

Click a word in the cloud or in the table.

locussymbolshort description

How it works

For every gene, the annotation texts are reduced to a set of concepts: lemmatised words (transporters → transporter), two-word phrases that co-occur far more than expected genome-wide (phosphate starvation, high affinity), UniProt keywords and, optionally, GO terms with their ancestors. A gene counts once per concept. Genes that are not annotated in the chosen concepts are left out of both the list and the background. For a concept carried by k of your n genes and K of the N background genes, P is the hypergeometric probability of seeing at least k. P values are corrected (Benjamini–Hochberg, or Benjamini–Yekutieli) over every concept that is testable on this background (carried by at least the minimum number of genes, and by at most 25% of the background), whether or not it occurs in your list; concepts carried by fewer genes of your list than the minimum enter with P = 1. All probabilities are computed in log scale, so very small values (below 10⁻³⁰⁰) stay exact. Concepts carried by the same genes are merged into one word (the most significant); a lone word is shown with the phrase that at least 90% of its genes share. Word size grows linearly with −log10 FDR (or log2 fold enrichment), colour = fold enrichment; grey words are nominal trends (uncorrected P ≤ 0.01), leads rather than findings. The same computation is implemented in the Python package, and the two give identical tables.