The datasets
Every definition, example, pronunciation, relation and translation stored in the app carries the identifier of the dataset it came from. The app shows that attribution on the entry itself, so you can always see where a particular line came from — not just which sources the app uses in general.
English Wiktionary
source_id: wiktionary
CC BY-SA 4.0 GFDL 1.3
Definitions and senses, grammatical labels, inflected forms, hyphenation, IPA for US and UK, etymologies, usage examples, sense-linked synonyms and antonyms, and translations into the target languages.
Used through the machine-readable extraction published by kaikki.org (wiktextract). Wiktionary text is dual-licensed under the Creative Commons Attribution-ShareAlike 4.0 International Licence and the GNU Free Documentation License 1.3. Contributed by the volunteer editors of en.wiktionary.org.
Open English WordNet 2024
source_id: oewn
CC BY 4.0
Thesaurus relations — synonyms, antonyms, hypernyms, hyponyms and similar-to links — and the sense inventory used to align related words to the right sense. WordNet glosses are used as a mapping aid and, where Wiktionary records no definition, as a fallback definition.
From Open English WordNet, licensed under the Creative Commons Attribution 4.0 International Licence.
CMU Pronouncing Dictionary
source_id: cmudict
BSD 2-Clause
ARPAbet phoneme sequences with stress marks for North American English. These are the primary input to the syllabification, IPA rendering and both respelling styles.
The CMU Pronouncing Dictionary is maintained by the CMUSphinx project at Carnegie Mellon University and released under the BSD 2-Clause Licence.
Tatoeba
source_id: tatoeba
CC BY 2.0 FR
Example sentences with human-written translations, used for the bilingual examples in entries and in the Translate tab.
Sentences from the Tatoeba Project are licensed under the Creative Commons Attribution 2.0 France Licence and remain the work of their individual contributors. Sentences whose contributors marked them as carrying a different licence are excluded from our build.
wordfreq
source_id: wordfreq
MIT
Word frequency estimates, used to decide which words are worth including and to rank suggestions and examples. No wordfreq text is displayed in the app; it informs ordering only.
The wordfreq library is released under the MIT Licence; its data is compiled from open corpora.
Dictionary Genie editorial
source_id: editorial
© Omnia Data Analytics LLC
Our own material: the notes attached to each Word of the Day, the Latin phrases and their glosses, and hand-verified pronunciation overrides where an automated result was wrong.
This is the only content in the app that we wrote. It is not offered under an open licence.
Share-alike, and what we owe back
Wiktionary is licensed under CC BY-SA 4.0, a share-alike licence. Our dictionary is a derivative of it, so the share-alike obligation follows through into the data we build.
We honour that in three ways:
- Attribution in place. Attribution is not buried in a settings screen. Every row of content in the app records its source, and the entry you are reading names the source it came from.
- The same licence. The derived dictionary data that ships inside Dictionary Genie is made available under the Creative Commons Attribution-ShareAlike 4.0 International Licence, the same terms we received it under.
- Corrections upstream. When someone reports a content error that also exists in the source, we fix it at the source as well, so the correction benefits everyone rather than only our users.
The application itself — its source code, interface, artwork, name and the editorial content listed above — is not covered by CC BY-SA and remains the property of Omnia Data Analytics LLC. The share-alike obligation attaches to the derived lexical data, not to the software that reads it.
Getting the derived data
The dictionary ships as a single content pack built by our open-data pipeline. Each release is accompanied by a build manifest recording the pack identifier and version, the schema version, the SHA-256 checksum and size of the pack, row counts, the languages included, and the frequency threshold used to select the corpus.
From the release of version 1.0, that manifest is published at:
https://dictionarygenie.app/sources/pack-manifest.json
If you want the derived dictionary data itself under CC BY-SA 4.0, email alend@omniadataanalytics.com naming the pack version from the manifest, and we will arrange a copy. We would also rather you took the original datasets above and built on those directly — they are better maintained than any snapshot we could hand you.
How the data is used
The pipeline that builds the pack does not paraphrase, summarise or invent. It parses, filters, normalises and joins. Specifically:
Respellings are the one thing in the dictionary that is computed rather than quoted. They are generated deterministically from recorded phonemes by a documented style guide, checked against a hand-verified gold set before any pack is accepted, and overridden by hand where the automated result is wrong.
Trademarks
Apple, iPhone, iPad, App Store, Siri and Spotlight are trademarks of Apple Inc., registered in the U.S. and other countries. Wiktionary and Wikimedia are trademarks of the Wikimedia Foundation, which does not endorse this app. Tatoeba, Open English WordNet, CMUSphinx and wordfreq are the projects of their respective maintainers, none of whom endorse this app. Use of these names here is attribution, as their licences require, and nothing more.
If you maintain one of these datasets and something on this page is wrong, incomplete, or not what your licence asks for, write to alend@omniadataanalytics.com and we will correct it.