Hmm, for now, after looking into this a bit:
I think the dataset side may actually be the more important long-term part of this project than the first model.
If I were choosing a default path from where you are now, I would not start by making the metadata much more academic, or by immediately training a larger model. I would mostly try to make the growing Hmar data traceable, reproducible, and easy to derive new datasets from.
A fairly small structure would already get you most of the benefit:
source/archive
↓
clean/derived documents
↓
sentences / paragraphs / task-specific datasets
↓
frozen evaluation snapshots
For the ML-facing datasets, even something as small as this would be useful:
{
"text": "...",
"source_id": "source_0042",
"doc_id": "doc_0042_01",
"register": "community_commentary"
}
Then the more verbose information — title, URL, date, rights/status, OCR notes, normalization notes, etc. — can live once in a source manifest rather than being repeated on every sentence.
Your corpus-archive already seems to be moving in roughly this direction, so I would probably build outward from that rather than redesigning everything.
The other thing I would start recording now, while the project is still manageable, is the exact dataset revision used for each model or benchmark. Your datasets are changing quickly enough that a model card saying “trained on sentences” will eventually become ambiguous. Hugging Face Datasets supports loading an exact branch/tag/commit through revision=, so something as simple as:
dataset = load_dataset(
"hmar-heritage-org/sentences",
revision="<commit-sha>",
)
makes an old result much easier to reproduce later.
So my rough priority order would be:
- Keep the source/archive separate from ML-ready derivatives.
- Give derived rows a lightweight path back to the source/document.
- Record dataset/model/tokenizer revisions for experiments.
- Write down how train/test was split.
- Freeze one small evaluation set and reuse it across future model versions.
- Only after that, worry about more elaborate schemas or more training.
And I would not worry too much about making the README sound “academic”. A useful dataset card can be very plain:
- What is this?
- Where did the text come from?
- What processing was done?
- What are the splits?
- What should users not assume about it?
- What version/revision is this?
- What are the reuse/rights conditions, if known?
That is enough to make the resource much more useful. HF’s own Dataset Card documentation describes the card mainly as documentation for understanding contents and intended use, and the Datasets guide explicitly says the complete set of fields is not required and even allows More Information Needed where something is not known yet.
So I would personally prefer a short README written in your own voice over a formally complete README that nobody in the project wants to maintain.
As for the original HmarBERT: I don’t think it is meaningless. I would treat it as a useful proof-of-concept for related-language transfer, rather than needing to decide whether it is already a definitive Hmar model.
I also took a small look at the released artifacts. The published MizBERT and HmarBERT vocabularies differ at exactly 303 token IDs, consistent with the model card’s description of 303 in-place cognate substitutions and zero collisions. Comparing the released checkpoints also suggests that the transformer backbone stayed relatively close to MizBERT while the embedding side moved substantially more. That is at least a plausible shape for a conservative related-language adaptation.
What that does not establish is that the 303 substitutions themselves caused the reported improvement. Continued Hmar training, vocabulary substitution, masking curriculum, and other training effects are mixed together unless an intermediate checkpoint/control is available.
That distinction is probably more useful than trying to decide whether the experiment was “right” or “wrong”.
Why starting from Mizo is not a strange idea
There is a broader family of work around exactly this general problem: how to reuse a pretrained model when the target language has little data and the source/target vocabularies do not line up perfectly.
For example:
These are not the same method as HmarBERT, and the language relationships/amounts of data are different, so I would not use them as evidence that the particular 303-word scheme is optimal.
But they do place the basic intuition —
preserve useful pretrained parameters, adapt the lexical interface, then continue training on the low-resource language
— inside a fairly recognizable research direction.
That is why I would not view “starting from MizBERT” as the obvious problem here.
In fact, because Hmar and Mizo are related, it is arguably a much more interesting baseline than starting from an unrelated BERT and simply hoping continued pretraining repairs everything.
The useful scientific question is then narrower:
What did the related-language initialization buy you, compared with the alternatives?
That can eventually be tested, but you do not need to answer it before the datasets themselves become useful.
A dataset layout that can stay simple
I would separate at least three conceptual layers.
1. Source/archive layer
This is where preservation and provenance live.
Something approximately like:
{
"source_id": "source_0042",
"title": "Example title",
"language": "hmr",
"source_type": "book",
"url": null,
"date": null,
"rights_status": "unknown",
"notes": "scanned copy; OCR corrected manually"
}
This does not have to become a grand universal ontology. It just needs to answer enough questions that someone can later understand where a derived row came from.
If some sources have different reuse conditions, keeping rights/status at source level is especially useful. I would avoid trying to make one repository-level license communicate more than is actually known about every underlying source.
That is a bookkeeping suggestion rather than a legal conclusion.
2. Derived corpus layer
Then sentences, paragraphs, etc. can stay pleasantly boring:
{
"text": "...",
"source_id": "source_0042",
"doc_id": "doc_0042_01",
"register": "literary_prose"
}
If useful later, you can add fields such as:
normalization_version
language_id_score
segment_id
but I would not add them until they solve an actual problem.
The important property is:
a derived row can be traced backward.
That gives you a lot later:
- remove or repair one problematic source;
- rebuild a dataset after improving OCR/normalization;
- create source-disjoint train/test splits;
- measure register balance;
- find accidental duplication;
- change filtering rules without losing the original material.
I would also preserve the source-faithful text somewhere rather than making normalized text the only surviving form. Orthography/normalization policy may change as the project learns more.
3. Frozen evaluation layer
Training datasets can keep growing.
Evaluation data should move more slowly.
For example:
eval/
hmar_mlm_v1.jsonl
manifest.json
with a small manifest such as:
{
"dataset_revision": "...",
"split_policy": "...",
"mask_selection": "...",
"seed": 1234,
"normalization": "...",
"scoring": "..."
}
Then HmarBERT-mini, mini-v2, or a future model can all be run against the same snapshot.
That creates a stable ruler while the underlying corpus continues growing.
Train/test splitting: I would document the unit, not assume a problem
I cannot tell from the public rows whether the current train/test division is already source/document-disjoint.
So I would not say there is leakage.
I would just make the split unit explicit.
There is an important difference between:
random sentence split
and:
hold out complete source documents
If several sentences from the same book/article appear on both sides, the test still measures something, but it is a somewhat easier question:
Can the model predict unseen passages drawn from sources it has partly seen?
A document/source holdout asks a stronger question:
Does it generalize to unseen source material?
Neither one has to replace the other. They can be two evaluation sets.
A simple decision tree would be:
Do you already know source_id/doc_id?
│
├─ yes
│ ├─ train/test already grouped by source/doc
│ │ └─ great — just document that
│ │
│ └─ rows were randomly split
│ └─ keep the existing split if useful,
│ and add a grouped holdout later
│
└─ no
└─ start recording IDs for new material;
don't block the whole project waiting for a perfect historical backfill
This matters because duplicate and near-duplicate material can make evaluation less independent. A useful general reference is Lee et al. 2022, Deduplicating Training Data Makes Language Models Better, which includes train/test overlap as one consequence of duplication.
But I would keep two concepts separate:
- deduplication: “is the same/similar text repeated?”
- grouped splitting: “did related passages from the same underlying source cross the boundary?”
You can deduplicate perfectly and still have both halves of the same book on opposite sides of a random sentence split.
For evaluation, I would reuse what you already have before inventing a bigger benchmark
The existing HmarBERT-mini benchmark material is potentially more valuable as a versioned, repeatable snapshot than as a one-off score report.
My next model-side experiment would probably be very boring:
Freeze the same examples + same masking rule + same scorer, and run HmarBERT-mini-v2 against exactly that.
That comparison can answer more than another training run, because tokenizer/model/data changes otherwise become entangled.
For future runs I would record:
model repo + commit
tokenizer repo + commit
dataset repo + commit
test row/source IDs
split construction rule
mask-selection rule
random seed
normalization used by scoring
metric definition
The HF Model Card documentation specifically treats training information, datasets, and evaluation results as useful model-card information, so this can live there without turning the repository into a paper.
One small metric note: BERT-style masked language models do not have left-to-right perplexity in quite the same sense as autoregressive models. There are established alternatives such as pseudo-log-likelihood / pseudo-perplexity; Salazar et al. 2020 is a useful reference.
I would not necessarily change your current numbers just because of terminology. The important practical step is to preserve exactly how the loss/PPL-like number was computed, so later versions can reproduce the same comparison.
A small note about the original 303-token experiment
This is probably a historical footnote rather than something I would redesign around now, but it was interesting.
Comparing the released MizBERT and HmarBERT vocabularies gives exactly:
303 changed token IDs
0 ID collisions
So the published description is unusually easy to audit.
There is one tokenizer-level wrinkle: under the released HmarBERT tokenizer, 267 of the 303 replacement strings tokenize directly back to their intended single swapped ID, while 36 do not. Those 36 are mostly forms involving diacritics/normalization behavior.
I would not conclude from that that those 36 rows were “unused” or that they damaged the model. BERT’s input embeddings and MLM output vocabulary are tied in the usual configuration, and continued training can move vocabulary rows through the output side as well.
In fact, comparing the released endpoint checkpoints showed those 36 rows moved quite strongly.
The more interesting broad observation was this:
MizBERT -> HmarBERT endpoint change
encoder-layer relative L2 drift: roughly 0.07–0.09
MLM-head drift: roughly 0.07
embedding-table drift: roughly 0.28
I would read that only as a descriptive checkpoint comparison:
a lot of adaptation happened around the lexical/embedding interface while much of the pretrained backbone remained comparatively close.
That seems perfectly compatible with the idea of conservative related-language adaptation.
It does not isolate whether the 303 substitutions caused the improvement. To answer that cleanly, you would want something like the Stage-1 pre-swap checkpoint evaluated under the same conditions.
If that checkpoint still exists, it is a nice optional ablation.
If it does not exist, I would not retrain a 110M model just to answer this Forum question. There are much higher-value things to do with the data.
Also, since the project has already moved on to newer Hmar-specific model/tokenizer experiments, I would treat this as a lesson from the first POC, not as a criticism of the current direction.
The compact decision tree I would use from here
1. Do derived rows have source/doc identity?
│
├─ YES
│ └─ expose/use it for manifests and grouped evaluation
│
└─ NO
└─ start with new material; backfill old data opportunistically
2. Is the train/test split source/document-disjoint?
│
├─ YES
│ └─ document the policy
│
├─ NO
│ └─ existing split can remain useful;
│ add one grouped holdout for stronger generalization testing
│
└─ UNKNOWN
└─ don't call it leakage;
record the construction rule on the next build
3. Is there already a frozen benchmark?
│
├─ YES
│ └─ run future models on exactly the same snapshot first
│
└─ NO
└─ freeze one small version before creating something huge
4. Does the old pre-swap/Stage-1 HmarBERT checkpoint still exist?
│
├─ YES
│ └─ optional: compare it with Stage 2/3 to separate
│ continued adaptation from vocabulary swapping
│
└─ NO
└─ no need to retrain just for the ablation
5. Does some metadata field actually support a decision?
│
├─ YES
│ └─ keep it
│
└─ NO
└─ probably don't add it yet
That last one may be the most important rule. 
Low-resource projects can easily accumulate a lot of infrastructure copied from much larger projects. I would optimize for information you will actually use, not schema completeness.
If I had to reduce all of this to one recommendation:
Treat the preserved Hmar sources as the long-lived asset, treat training datasets as reproducible views of those sources, and treat benchmarks as frozen/versioned views.
Then the models can keep changing without losing the history of what they were trained and tested on.
And for what it is worth, the original HmarBERT experiment looks useful to me precisely because it does not need to be the final answer. It gives you a related-language transfer baseline; the newer models give you from-scratch/Hmar-specific alternatives; and a stable data/evaluation layer would make those choices increasingly comparable over time.