by devmount
Toolkit to obtain and preprocess German text corpora, train models and evaluate them with generated testsets. Built with Gensim and Tensorflow.
# Add to your Claude Code skills
git clone https://github.com/devmount/GermanWordEmbeddingsGuides for using testing skills like GermanWordEmbeddings.
Last scanned: 10/8/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-10-08T11:10:01.254Z",
"npmAuditRan": true,
"pipAuditRan": false,
"promptInjectionRan": true
}GermanWordEmbeddings is an open-source testing skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by devmount. Toolkit to obtain and preprocess German text corpora, train models and evaluate them with generated testsets. Built with Gensim and Tensorflow. It has 245 GitHub stars.
Yes. GermanWordEmbeddings passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/devmount/GermanWordEmbeddings" and add it to your Claude Code skills directory (see the Installation section above).
GermanWordEmbeddings is primarily written in Jupyter Notebook. It is open-source under devmount on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other Testing skills you can browse and compare side by side. Open the Testing category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh GermanWordEmbeddings against similar tools.
No comments yet. Be the first to share your thoughts!
Top skills in this category by stars
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
There has been a lot of research about the training of word embeddings on English corpora. This toolkit applies gensim's word2vec on German corpora to train and evaluate German word embeddings. An overview about the project, evaluation results and download links can be found on the project's website or directly in this repository.
[!NOTE] This project is old (2015) and trained word embeddings on German text long before LLMs were called AI. I still actively maintain this repo, the models from the past still load and can be used.
Make sure you have Python 3.12 or 3.13 installed, as well as the required libraries and NLTK data, preferably in a virtual environment:
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m nltk.downloader punkt_tab stopwords
Now you can download word2vec_german.sh and execute it in your shell to automatically download this toolkit and the corresponding corpus files and do the model training and evaluation. The script expects the libraries and NLTK data from above to be installed and trains a model with 300 dimensions, a window size of 5, 10 negative samples and a minimum word count of 50. Be aware that this could take a huge amount of time!
You can also clone this repository and use my already trained model to play around with the evaluation and visualization.
If you just want to see how the different Python scripts work, have a look into the code directory to see Jupyter Notebook script output examples. Mind that these outputs are from the original runs in 2015.
There are multiple possibilities for obtaining huge German corpora that are publicly available and free to use:
wget https://dumps.wikimedia.org/dewiki/latest/dewiki-latest-pages-articles.xml.bz2
Shuffled German news of the years 2007 to 2013:
for i in 2007 2008 2009 2010 2011 2012 2013; do
wget https://www.statmt.org/wmt14/training-monolingual-news-crawl/news.$i.de.shuffled.gz
done
The models published with this toolkit are based on the German Wikipedia and only the German news of 2013.
This tool preprocesses the raw wikipedia XML corpus with the WikiExtractor (a Python Script from Giuseppe Attardi to filter a Wikipedia XML Dump, installed via requirements.txt) and some shell instructions to filter all XML tags and quotations:
python -m wikiextractor.WikiExtractor -c -b 25M -o extracted dewiki-latest-pages-articles.xml.bz2
find extracted -name '*bz2' \! -exec bzip2 -k -c -d {} \; > dewiki.xml
sed -i 's/<[^>]*>//g' dewiki.xml
sed -i 's|["'\''„“‚‘]||g' dewiki.xml
rm -rf extracted
The German news already contain one sentence per line and don't have any XML syntax overhead. Only quotation marks need to be removed:
for i in 2007 2008 2009 2010 2011 2012 2013; do
gzip -d news.$i.de.shuffled.gz
sed -i 's|["'\''„“‚‘]||g' news.$i.de.shuffled
done
Afterwards, the preprocessing.py script can be called on these corpus files with the following options:
| flag | default | description |
|---|---|---|
| -h, --help | - | show a help message and exit |
| -p, --punctuation | False | filter punctuation tokens |
| -s, --stopwords | False | filter stop word tokens |
| -u, --umlauts | False | replace German umlauts with their respective digraphs |
| -b, --bigram | False | detect and process common bigram phrases |
| -t [ ], --threads [ ] | NUMBER_OF_PROCESSORS | number of worker threads |
| --batch_size [ ] | 32 | batch size for sentence processing |
Example usage:
python preprocessing.py dewiki.xml corpus/dewiki.corpus -psub
for file in *.shuffled; do python preprocessing.py $file corpus/$file.corpus -psub; done
Mind that the -b flag creates an additional .bigram file next to each corpus file. As the training uses every file in the corpus directory, only keep one of both versions, e.g. with rm corpus/*.corpus.
Models are trained with the help of the training.py script with the following options:
| flag | default | description |
|---|---|---|
| -h, --help | - | show a help message and exit |
| -s [ ], --size [ ] | 100 | dimension of word vectors |
| -w [ ], --window [ ] | 5 | size of the sliding window |
| -m [ ], --mincount [ ] | 5 | minimum number of occurrences of a word to be considered |
| -t [ ], --threads [ ] | NUMBER_OF_PROCESSORS | number of worker threads to train the model |
| -g [ ], --sg [ ] | 1 | training algorithm: Skip-Gram (1), otherwise CBOW (0) |
| -i [ ], --hs [ ] | 1 | use of hierarchical softmax for training |
| -n [ ], --negative [ ] | 0 | use of negative sampling for training (usually between 5-20) |
| -o [ ], --cbowmean [ ] | 0 | for CBOW training algorithm: use sum (0) or mean (1) to merge context vectors |
| -f, --full | False | additionally store the full model in <target>.full to resume training later |
| -r [ ], --resume [ ] | - | full model file to resume training from with new corpora |
Example usage:
python training.py corpus/ my.model -s 200 -w 5
Mind that the first parameter is a directory and that every contained file will be taken as a corpus file for training.
If the time needed to train the model should be measured and stored into the results file, this would be a possible command:
{ time python training.py corpus/ my.model -s 200 -w 5; } 2>> my.model.result
By default only the word vectors are stored, which is sufficient for evaluation and visualization. To be able to continue the training later, e.g. with new words, the full model has to be stored with the -f flag. It can then be passed to the -r option together with a directory that only contains the additional corpus files:
python training.py corpus/ my.model -s 200 -w 5 -f
python training.py corpus_new/ my.model -r my.model.full -f
Mind that the model parameters (-s, -w, -m, -g, -i, -n, -o) are taken from the resumed model and are ignored when given, and that a full model needs considerably more disk space than the word vectors alone.
To compute the vocabulary of a trained model, the vocabulary.py script can be used:
python vocabulary.py my.model my.model.vocab
Mind that the binary model format doesn't contain word frequencies, so the stored counts only reflect the frequency rank of each word.
To create test sets and evaluate trained models, the evaluation.py script can be used. It's possible to evaluate both syntactic and semantic features of a trained model. For a successful creation of testsets, the following source files should be created before starting the script (see the configuration part in the script for more information).
With the syntactic test, features like singular, plural, 3rd person, past tense, comparative or superlative can be evaluated. Therefore there are 3 source files: adjectives, nouns and verbs. Every file contains a unique word with its conjugations per line, divided by a dash. These combination patterns can be entered in the PATTERN_SYN constant in the script configuration. The script now combines each word with 5 random other words according to the given pattern, to create appropriate analogy questions. Once the data file with the questions is created, it can be evaluated. Normally the evaluation can be done by gensim's word analogy evaluation function, but to get a more specific evaluation result (correct matches, top n matches and coverage), this project uses its own accuracy functions (test_most_similar_groups() and test_most_similar() in evaluation.py).
The given source files of this project contain 100 unique nouns with 2 patterns, 100 unique adjectives with 6 patterns and 100 unique verbs with 12 patterns, resulting in 10k analogy questions. Here are some examples for possible source files:
Possible pattern: basic-comparative-superlative
Example content:
gut-besser-beste
laut-lauter-lauteste
Possible pattern: singular-plural
Example content:
Bild-Bilder
Name-Namen
See src/nouns.txt
Possible pattern: basic-1stPersonSingularPresent-2ndPersonPluralPresent-3rdPersonSingularPast-3rdPersonPluralPast
Example content:
finden-finde-f