# Data sources & attribution

`data/jlpt.json` (and any vector files generated from it, e.g. `data/vectors.int8.bin`) are
**adapted from [OpenJLPT](https://github.com/evanclan/OpenJLPT)** (vocabulary + kanji) and
**[AnchorI/jlpt-kanji-dictionary](https://github.com/AnchorI/jlpt-kanji-dictionary)** (kanji extras),
and are distributed under
**[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/)** (full text: `LICENSE-CC-BY-SA-4.0.txt`).

Required credit line (also shown in the page footer):

> Contains data from [OpenJLPT](https://github.com/evanclan/OpenJLPT) (CC BY-SA 4.0), which uses
> JMdict and KANJIDIC2 (EDRDG), Jonathan Waller's JLPT lists, and Tatoeba.
> Kanji details include [AnchorI/jlpt-kanji-dictionary](https://github.com/AnchorI/jlpt-kanji-dictionary) (MIT).

## Upstream sources (as stated in OpenJLPT's NOTICE.md)

| Source | Used for | License |
|---|---|---|
| Jonathan Waller's JLPT Resources (https://www.tanos.co.uk/jlpt/) | N5–N1 level assignments and English glosses | CC BY |
| JMdict, EDRDG (https://www.edrdg.org/jmdict/j_jmdict.html) | `jmdict_id`, part of speech, verification/repair of readings, spellings and glosses | CC BY-SA 4.0 (see https://www.edrdg.org/edrdg/licence.html) |
| Tatoeba (https://tatoeba.org) | Example sentences (Japanese + English); each carries its Tatoeba sentence ID | CC BY 2.0 FR |
| AnchorI/jlpt-kanji-dictionary (https://github.com/AnchorI/jlpt-kanji-dictionary) | Kanji strokes / radical / frequency / description | MIT |
| ECDICT (https://github.com/skywind3000/ECDICT) | English vocab sets: IELTS / TOEFL / GRE / CET-4 / CET-6 / 考研 / 高考 / 中考 — word, phonetic, Chinese translation | MIT |

## Changes made in this project

- Kept only these fields per entry: id, word, reading, romaji, meanings, level, part of speech,
  and at most 2 example sentences (furigana markup removed).
- Vectors (if generated) are derived from the same text and are therefore also CC BY-SA 4.0.

## Obligations to keep in mind

1. Keep the credit line and a link to the CC BY-SA 4.0 license wherever the data is shown.
2. Distribute adapted data (JSON, vectors, databases built from it) under CC BY-SA 4.0.
3. JLPT levels are an unofficial community list, not official test content.
4. EDRDG asks redistributors of JMdict-derived data to keep it reasonably current: re-run
   `python3 tools/build-data.py` (and rebuild vectors) periodically.

This file is a practical summary, not legal advice; check the upstream licenses yourself.
