{"id":653,"date":"2026-09-15T16:10:37","date_gmt":"2026-09-15T07:10:37","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/eleutherai-releases-bergson-leaderboard-baseline-gpt2-model\/"},"modified":"2026-09-18T21:42:02","modified_gmt":"2026-09-18T12:42:02","slug":"eleutherai-releases-bergson-leaderboard-baseline-gpt2-model","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/eleutherai-releases-bergson-leaderboard-baseline-gpt2-model\/","title":{"rendered":"EleutherAI Releases Bergson Leaderboard Baseline GPT-2 Model"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/EleutherAI\/bergson-wikitext-gpt2-leaderboard\">EleutherAI\/bergson-wikitext-gpt2-leaderboard<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-15<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>safetensors<\/td>\n<\/tr>\n<tr>\n<td>Paper<\/td>\n<td><a href=\"https:\/\/arxiv.org\/abs\/2303.14186\">arXiv:2303.14186<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>EleutherAI has released &#8220;EleutherAI\/bergson-wikitext-gpt2-leaderboard&#8221;, the baseline model for the training data attribution benchmark &#8220;bergson leaderboard&#8221;.<\/p>\n<p>This model is based on openai-community\/gpt2 and has been fine-tuned for 4 epochs using 4,608 chunks of the WikiText dataset (EleutherAI\/bergson-wikitext-512-chunks). According to the model card, the held-out loss has decreased from 3.545 to 3.111.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Base model: openai-community\/gpt2<\/li>\n<li>License: apache-2.0<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Publishing model card details, scores from the bergson leaderboard evaluating various data attribution methods are disclosed. The table provided is as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Method<\/th>\n<th style=\"text-align: center;\">Proponent QLD [95% CI]<\/th>\n<th style=\"text-align: center;\">LDS [95% CI]<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\">MAGIC<\/td>\n<td style=\"text-align: center;\">0.100 [0.090, 0.112]<\/td>\n<td style=\"text-align: center;\">0.931 [0.925, 0.936]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Eigenvalue-corrected Shampoo<\/td>\n<td style=\"text-align: center;\">0.071 [0.060, 0.082]<\/td>\n<td style=\"text-align: center;\">0.517 [0.491, 0.539]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">EK-FAC<\/td>\n<td style=\"text-align: center;\">0.070 [0.058, 0.082]<\/td>\n<td style=\"text-align: center;\">0.454 [0.426, 0.479]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">KFAC<\/td>\n<td style=\"text-align: center;\">0.067 [0.056, 0.080]<\/td>\n<td style=\"text-align: center;\">0.420 [0.391, 0.446]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">BM25<\/td>\n<td style=\"text-align: center;\">0.062 [0.048, 0.076]<\/td>\n<td style=\"text-align: center;\">0.220 [0.185, 0.252]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><a href=\"https:\/\/huggingface.co\/spaces\/mteb\/leaderboard\">Qwen3-Embedding-8B<\/a> semantic search<\/td>\n<td style=\"text-align: center;\">0.049 [0.038, 0.061]<\/td>\n<td style=\"text-align: center;\">0.132 [0.093, 0.169]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">TrackStar (no optimizer correction, projection 64)<\/td>\n<td style=\"text-align: center;\">0.045 [0.036, 0.055]<\/td>\n<td style=\"text-align: center;\">0.270 [0.240, 0.295]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">TRAK (8-model ensemble)<\/td>\n<td style=\"text-align: center;\">0.032 [0.024, 0.040]<\/td>\n<td style=\"text-align: center;\">0.138 [0.111, 0.165]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">SOURCE (Adam)<\/td>\n<td style=\"text-align: center;\">0.024 [0.018, 0.030]<\/td>\n<td style=\"text-align: center;\">0.154 [0.126, 0.181]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Gradient cosine similarity<\/td>\n<td style=\"text-align: center;\">0.021 [0.016, 0.027]<\/td>\n<td style=\"text-align: center;\">0.156 [0.131, 0.181]<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Activation similarity<\/td>\n<td style=\"text-align: center;\">0.000 [-0.000, 0.001]<\/td>\n<td style=\"text-align: center;\">0.110 [0.070, 0.149]<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>According to the explanations of the metrics provided by the publishers, LDS (Linear datamodeling score) indicates the accuracy of global data ranking methods based on influence, while QLD (Query loss difference) shows how much the model loss increases for held-out queries when top 1% influence data is removed and retrained compared to a random exclusion baseline.<\/p>\n<p>Based on the scores in the table, MAGIC recorded 0.100 [0.090, 0.112] in Proponent QLD and 0.931 [0.925, 0.936] in LDS, representing the highest values among all methods. On the other hand, Activation similarity yielded the lowest results with 0.000 [-0.000, 0.001] for Proponent QLD and 0.110 [0.070, 0.149] for LDS. Additionally, the Proponent QLD scores for EK-FAC and Eigenvalue-corrected Shampoo were 0.070 and 0.071 respectively, showing a very close margin.<\/p>\n<p>Regarding the file structure included in the repository, the model card provides the following correspondence table:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>path<\/th>\n<th>what it is<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>exported\/checkpoint-72\/<\/code><\/td>\n<td>the final model, scored by every leaderboard method<\/td>\n<\/tr>\n<tr>\n<td><code>exported\/checkpoint-{12,24,36,48,60}\/<\/code><\/td>\n<td>interval checkpoints with <code>optimizer.pt<\/code>, used by SOURCE and the checkpoint-averaged variants<\/td>\n<\/tr>\n<tr>\n<td><code>optimizer.pt<\/code>, <code>config.yaml<\/code><\/td>\n<td>final AdamW second moments (TrackStar-Adam) and the training config<\/td>\n<\/tr>\n<tr>\n<td><code>trak_ensemble\/s{0..3}_subset_{0,1}\/<\/code><\/td>\n<td>the eight GPT-2 models trained on independent random 50% subsets for the TRAK row (<a href=\"https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext\/trak_ensemble\"><code>trak_ensemble\/train_s*.yaml<\/code><\/a>)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Strengths and Use Cases<\/h2>\n<p>According to the model card and metadata, this model is intended for use cases related to training-data-attribution, influence-functions, and model interpretability.<\/p>\n<p>Rather than serving as a general-purpose conversational or coding model for specific tasks, it is designed as a benchmark model to measure and compare the impact that various data attribution methods have on the model within the <a href=\"https:\/\/bergson.readthedocs.io\/en\/latest\/leaderboard.html\">bergson leaderboard<\/a>. The retraining bank and set of score results are provided at <a href=\"https:\/\/huggingface.co\/datasets\/EleutherAI\/bergson-wikitext-gpt2-leaderboard-bank\"><code>EleutherAI\/bergson-wikitext-gpt2-leaderboard-bank<\/code><\/a>.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>8GB (RTX 4060 \/ 3060 Ti, etc.)<\/td>\n<td>\u305d\u306e\u307e\u307e\u306e\u7cbe\u5ea6<\/td>\n<td>6.5GB<\/td>\n<td>7.8GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is available for download from Hugging Face. The distribution format includes safetensors, and since the repository is not gated, it can be acquired without any prior application or agreement procedures.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/eleutherai-olmo-3-7b-reward-hacking-models\/\">EleutherAI Releases OLMo-3-7B Models for Reward-Hacking Research<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/EleutherAI\/bergson-wikitext-gpt2-leaderboard\">https:\/\/huggingface.co\/EleutherAI\/bergson-wikitext-gpt2-leaderboard<\/a><\/li>\n<li><a href=\"https:\/\/bergson.readthedocs.io\/en\/latest\/leaderboard.html\">https:\/\/bergson.readthedocs.io\/en\/latest\/leaderboard.html<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/datasets\/EleutherAI\/bergson-wikitext-512-chunks\">https:\/\/huggingface.co\/datasets\/EleutherAI\/bergson-wikitext-512-chunks<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext\/2_interval.yaml\">https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext\/2_interval.yaml<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/EleutherAI\/bergson\">https:\/\/github.com\/EleutherAI\/bergson<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/datasets\/EleutherAI\/bergson-wikitext-gpt2-leaderboard-bank\">https:\/\/huggingface.co\/datasets\/EleutherAI\/bergson-wikitext-gpt2-leaderboard-bank<\/a><\/li>\n<li><a href=\"https:\/\/arxiv.org\/abs\/2303.14186\">https:\/\/arxiv.org\/abs\/2303.14186<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext\">https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/collections\/EleutherAI\/data-attribution-6a8d013ccc372b7fd6abce3e\">https:\/\/huggingface.co\/collections\/EleutherAI\/data-attribution-6a8d013ccc372b7fd6abce3e<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/spaces\/mteb\/leaderboard\">https:\/\/huggingface.co\/spaces\/mteb\/leaderboard<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext\/trak_ensemble\">https:\/\/github.com\/EleutherAI\/bergson\/tree\/main\/examples\/compare_wikitext\/trak_ensemble<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>EleutherAI has released the baseline model EleutherAI\/bergson-wikitext-gpt2-leaderboard for the training data attribution benchmark bergson leaderboard.<\/p>\n","protected":false},"author":1,"featured_media":658,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[1233,373,1235,1237,524,117],"class_list":["post-653","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-bergson-en","tag-eleutherai-en","tag-gpt-2-en","tag--en"],"lang":"en","translations":{"en":653,"ja":651},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/653","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=653"}],"version-history":[{"count":6,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/653\/revisions"}],"predecessor-version":[{"id":1670,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/653\/revisions\/1670"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/658"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=653"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=653"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=653"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}