Vocab Bloom Hub
هذه الصفحة متاحة بالإنجليزية فقط.

Moving a dictionary between instances offline

The dictionary import normally downloads the published dataset from HuggingFace. Installations without internet access — corporate networks, air-gapped labs, CI — and admins who edited an exported dataset can load the same files from a local source instead:

  • files uploaded through the admin UI — the zip that Export dictionary produces (the Archive tab), or the dataset files in their own slots with the manifest either as a file or typed by hand (the Separate files tab);
  • a directory or zip on the server — inside the folder named by DICTIONARY_IMPORT_DIR (a mounted volume, for instance).

Both go through the same import pipeline as the HuggingFace source; only the download stage is skipped.

Dataset format

A dataset is the set of files the export writes, flat or inside a single wrapper folder:

manifest.json
vocab-bloom-hub-en-words.jsonl
vocab-bloom-hub-en-phrasal-verbs.jsonl
vocab-bloom-hub-en-grammar-patterns.jsonl
vocab-bloom-hub-en-phrases.jsonl
vocab-bloom-hub-en-meanings.jsonl
vocab-bloom-hub-en-meaning-translations.<lang>.jsonl   # one per language: .ru, .es, .fr, .de, .pt, .zh, .ar
vocab-bloom-hub-en-short-translations.<lang>.jsonl

The entry files (words, phrases, grammar-patterns) carry the entries themselves and their forms; phrasal-verbs is the linking map of base verbs to their phrasal variants. The collections are files of their own, one line per row next to the key of its parent:

FileOne line perKey of the parent
meaningsmeaning (with its links)word, part_of_speech
meaning-translations.<lang>translation of a meaningword, part_of_speech, meaning_sort_order, meaning_title
short-translations.<lang>short translation of an entryword, part_of_speech

The translations are one file per language (<lang> is a member of available_languages), so a file stays a manageable size and a language loads on its own; a language without rows has no file. Exports made before this split wrote one combined meaning-translations.jsonl / short-translations.jsonl for every language — the import still reads those.

Phrases and grammar patterns are keyed the same way, with phrase / grammar_pattern as the part of speech. The lines follow the order of the entry files, then the natural keys of the rows, so two exports of the same data are byte-identical. Datasets published before this layout nest meanings and short_translations inside the entry lines; the import still reads them.

Only the .jsonl files you actually have are needed — at least one of them; every file but the words file is optional, and so is manifest.json. The collection files are imported after the entry files and go to the entries and meanings that exist by then, whether they came from this dataset or were there before: a file of translations alone — one language, say — loads into a dictionary that already holds the entries, and rows the dictionary already has (a meaning with the same sort order and title, a translation in the same language with the same title, a short translation in the same language with the same description) are skipped like duplicate entries. When the manifest is present its version is stored as Your version after the import (and its synonym / antonym link counts refine the progress bar); without it the version stays unknown. The license and attribution fields an export writes into the manifest (DATA_LICENSE.md) are carried along and not checked. Line counts for the progress bar are always taken from the files themselves, so a hand-assembled dataset needs no bookkeeping.

ملاحظة

The dataset is validated before the import starts and is rejected (dataset_invalid) when there is no .jsonl file at all, when manifest.json is malformed, or when the folder / archive contains any file with an unknown name. OS artefacts (.DS_Store, __MACOSX/) are ignored.

Export on A → copy → import on B

  1. On instance A open Managing → Export dictionary (or call GET /api/en/dictionary/export and download the archive from GET /api/en/dictionary/export/download/:exportId). You get vocab-bloom-hub-en-export.zip.

  2. Copy the zip to a machine that can reach instance B.

  3. On instance B open Managing → Import dictionary → Archive, drop the zip into the upload area and press Start importing. The progress stream is the same as for the HuggingFace import; the manifest version is stored as Your version when the import completes.

    To load the files separately open the Separate files tab instead: one slot per file (the slot decides what the file is, its own name does not matter), and the manifest either as manifest.json or filled in by hand (version, optional synonym / antonym link counts).

Equivalent API calls (the admin cookie or a Bearer token is required). The multipart fields are archive for the whole zip, or words, phrasal_verbs, grammar_patterns, phrases, meanings, meaning_translations_<lang>, short_translations_<lang> (one slot per language, e.g. short_translations_es) and manifest for the separate files; the text fields version, synonym_links and antonym_links stand in for (and override) a manifest file:

The examples use localhost:3010, the port of a start without Docker; under docker compose the API is on localhost:3240 by default.

# the whole archive
curl -N -b cookies.txt -F archive=@vocab-bloom-hub-en-export.zip \
  http://localhost:3010/api/en/dictionary/import/upload

# only the words, the version typed by hand
curl -N -b cookies.txt -F words=@my-words.jsonl -F version=1.4.0 \
  http://localhost:3010/api/en/dictionary/import/upload

# translations only, for entries the dictionary already has
curl -N -b cookies.txt -F short_translations_es=@es-short-translations.jsonl \
  -F meaning_translations_es=@es-meaning-translations.jsonl \
  http://localhost:3010/api/en/dictionary/import/upload

ملاحظة

Each upload is limited to 512 MiB; the uploaded files are deleted from the server once the import has finished, whether or not it succeeded.

Datasets on the server (DICTIONARY_IMPORT_DIR)

When the zip is already on the server — mounted into a container, copied by a deploy script — point DICTIONARY_IMPORT_DIR at the folder holding it:

DICTIONARY_IMPORT_DIR=/data/dictionary-imports

The Archive tab then also lists what the folder offers (zip archives and sub-folders containing a manifest.json, one level deep) next to the upload area, and the request names the pick relative to that folder:

curl -N -b cookies.txt -H 'Content-Type: application/json' \
  -d '{"source":{"kind":"file","path":"vocab-bloom-hub-en-export.zip"}}' \
  http://localhost:3010/api/en/dictionary/import

ملاحظة

Paths are resolved inside the import directory only: .., absolute paths and symlinks pointing elsewhere are rejected (dataset_file_not_found). Without the variable the request fails with import_dir_not_configured and the Archive tab offers uploads only. The server never modifies or deletes the files in that folder.

GET /api/en/dictionary/import/sources returns the same listing the UI shows:

{
  "import_dir_configured": true,
  "files": [
    {
      "path": "vocab-bloom-hub-en-export.zip",
      "kind": "zip",
      "size": 12345678,
      "modified_at": "2026-09-15T18:18:25.486Z"
    },
    { "path": "hand-assembled", "kind": "directory", "size": 0, "modified_at": "2026-09-15T18:18:25.486Z" }
  ],
  "revisions": ["v0.1.0"]
}

revisions are the version tags of the published dataset repository, newest first (the Dataset version selector of the HuggingFace tab; [] when the HuggingFace refs API is unreachable).

Notes

مهم

Importing does not wipe the database first: records that already exist (same word, part of speech and form) are skipped, the same way the HuggingFace import behaves. To replace existing entries with the dataset content (keeping the entries you edited), run the import in update mode — see operations.md.

نصيحة

With Docker Compose, mount the folder into the server container and set the variable to the mount point, e.g. - ./imports:/data/dictionary-imports and DICTIONARY_IMPORT_DIR=/data/dictionary-imports.

ملاحظة

pg_dump / pg_restore remain the right tool for moving a whole database including ids; the dataset route is for the dictionary content in its portable, diffable form.