Using Hugging Face to Extract Dates for a Date Duration Calculator

I’m experimenting with Hugging Face for a small NLP project involving date and time expressions.

The idea is to build a date-duration calculator where users don’t necessarily have to enter a date in a strict format. Instead, they could write something like:

We started dating on August 18, 2024.

The NLP layer would identify August 18, 2024 as a date and normalize it into something like:

{
  "event": "relationship_start",
  "date": "2024-08-18"
}

After that, the actual duration calculation would be handled separately by normal date/time logic. For example, a tool such as Kalkulator Jadian uses the idea of calculating elapsed relationship time from a starting date, so I’m interested in how an NLP model could make that input more flexible.

My main question is about the best Hugging Face approach for this pipeline.

Would you recommend using a token-classification/NER model to extract the date first and then passing the extracted value to a deterministic date parser? Or would it be better to fine-tune a model specifically for temporal expressions?

I’d also like to handle inputs such as:

  • August 18, 2024

  • 18/08/2024

  • last Friday

  • two weeks ago

  • our anniversary was on 18 August

I’m trying to keep the architecture simple:

Natural language
      ↓
Date extraction
      ↓
Date normalization
      ↓
Date validation
      ↓
Duration calculation
      ↓
Days / hours / minutes / seconds

The important part is keeping the NLP model responsible for understanding the user’s text while leaving the actual date arithmetic to deterministic code.

For someone learning Hugging Face, would NER/token classification be a reasonable starting point for this type of project, or is there a better model/task for temporal expression extraction and normalization?

Separating date extraction from normalization sounds like a good approach:


Short answer: NER is a reasonable starting point, but I would not fine-tune first

For a project intended to learn Hugging Face, I would start with a pretrained DATE-capable extractor (NER/token classification, or a model that accepts a date label), inspect the actual text spans it returns, and pass those spans to a separate date normalizer. This fits the architecture you described. I would keep the final duration arithmetic in ordinary code, as you proposed.

I would not make training a specialist temporal model the first prerequisite. Try an existing extractor and a parser baseline on your five examples before deciding that you need labeled training data. Importantly, “NER” and “temporal-expression model” are not mutually exclusive categories: some temporal taggers are themselves token classifiers, but trained for DATE, TIME, DURATION, and SET, rather than generic named entities. The HF token-classification guide and this temporal BERT model card illustrate the distinction.

Two starting routes make sense:

  • Learning HF extraction: Try a pretrained DATE extractor, such as GLiNER small-v2.1 with the label date, or a temporal token classifier. Feed the returned complete phrases into a conventional parser. This reveals missed or clipped spans without fine-tuning.
  • Finishing a small calculator quickly: Parse known formats directly; use a natural-language date parser for phrases such as two weeks ago; introduce an HF extractor only if users’ full sentences cause missing or fragmented spans. dateparser is a Python baseline and chrono-node is a JavaScript option. Neither handles every sentence or event relation.

I would add just one explicit boundary to your diagram: after normalization, decide which date refers to the event being calculated (relationship_start, another event, an alternative, or a negated date), and whether the input is sufficiently specified to accept. A parseable date is not always the intended event date.

Suggested default pipeline

User text + known language/date order + reference date/time
        |
        v
Extract candidate spans
  - known formats / parser search as a baseline
  - optionally, a pretrained HF DATE/temporal extractor
        |
        v
Normalize EACH COMPLETE candidate
  - strict ISO or known format -> strict calendar parser
  - relative words -> parser + fixed reference date/time
  - incomplete/ambiguous -> preserve uncertainty
        |
        v
Link candidate to event + consider negation/alternatives
        |
        v
Validate / accept / clarify / abstain
        |
        v
Calculate calendar days OR elapsed seconds with normal code
        |
        v
Display days / hours / minutes / seconds by an explicit rule

A clipped last Friday points to extraction; a misread 2024-08-18 to normalization; a breakup date selected as a start to event linking; a wrong seconds counter to time policy. Only the first is directly fixed by better DATE span tagging.

A minimal first test

Your five phrases are a good starting panel. I would add just two contrasts: a sentence with two competing dates and negation, and a sentence where a date belongs to a different event. For each, save the extracted substring, normalized value (or reason unresolved), event role, and whether confirmation is needed.

Input What to check
August 18, 2024 Full span; resolves to 2024-08-18
18/08/2024 Numeric order and locale policy
last Friday Preserve last; fix a reference date for testing
two weeks ago Relative base and whether you mean date or exact instant
our anniversary was on 18 August Month/day found; year is missing; anniversary is not automatically the relationship start
Our start date was 2024-08-18, not last Friday Find both expressions but select the non-negated start date
We booked tickets for August 18, 2024 Date found, but do not assign relationship_start

Add one no-date negative to measure false positives; no GPU is needed.

Why I would keep these stages separate: a small independent probe

I tried a small, hand-built English diagnostic—not your application or a representative benchmark:

  • For We started dating on Sept. 18, 2024, our long-text dateparser search fragmented the date while a pretrained GLiNER date path found the complete Sept. 18, 2024 span. This is a concrete reason to try a DATE extractor when free-text search loses spans.
  • We also had cases where the extractor found the correct ISO YYYY-MM-DD substring but our shared parser settings interpreted it incorrectly. Sending complete ISO dates through a strict ISO branch improved the final result without retraining or rerunning the extractor. The error was in our pipeline configuration, not proof that dateparser generally cannot handle ISO dates.

In another case, a naive hybrid selected yesterday inside the day before yesterday. Adding more candidate generators can create more choices—and a worse answer—unless the candidate-selection step knows about complete spans, overlapping phrases, negation, and event roles.

Default: try pretrained DATE extraction, keep strict/relative parsing as a baseline, and fine-tune only if representative span-extraction failures remain.

Which model or component should come first?

NER, temporal tagging, and parsing solve related but different tasks

A token classifier labels tokens or spans in text. But not every checkpoint marketed as NER knows the label DATE. Some generic checkpoints expose only PER, ORG, LOC, or MISC; always inspect the model’s id2label/label scheme or model card. Conversely, a time-oriented token classifier may label DATE, TIME, DURATION, and recurring SET expressions yet still require another component to turn them into an actual calendar date.

The temporal BERT checkpoint is a useful example. Its author explicitly notes that the accompanying repository’s postprocessing matters for deciphering the model output. Running a generic pipeline("token-classification") and treating those raw tags as the entire published system would not be a fair evaluation of that approach.

GLiNER is a different useful experiment: you can supply the desired label at inference time. That is attractive for learning what the model’s span extraction does without immediately preparing a training set. It is not automatically a date normalizer, event linker, or guarantee of correct dates.

A small Hugging Face experiment

# Example extraction-only step; install/test a compatible GLiNER version.
from gliner import GLiNER

extractor = GLiNER.from_pretrained("urchade/gliner_small-v2.1")
text = "We started dating on Sept. 18, 2024."
items = extractor.predict_entities(text, ["date"], threshold=0.35)
for item in items:
    print(item["text"], item.get("start"), item.get("end"), item.get("score"))

Treat threshold=0.35 as a test setting, not a calibrated probability that a date is correct. Keep the exact span and offsets. Test date against date/time/duration only if those other expression types matter to the app; adding labels may change candidate boundaries and count. Pin the model, library and runtime versions for a serious comparison.

A parser-first experiment

If inputs tend to be short date phrases, beginning with known-format parsing may be enough. A calendar date such as 2024-08-18 does not need an NLP model to be recognized and validated. A relative phrase such as two weeks ago does need a reference date/time and a clear meaning. For a long sentence, dateparser.search.search_dates can be a baseline, but its documentation warns about limitations of search in longer text. For JavaScript, chrono-node can return matched text, offsets, components, and reference-date-aware results.

Between a parser and a transformer, there is a useful middle option: a small, inspectable pattern recognizer for explicit dates or well-known expressions, perhaps alongside an spaCy EntityRuler. If the expected input vocabulary is narrow, that may be the least expensive reliable component.

When fine-tuning becomes justified

Fine-tune when you have representative examples showing a repeated span-extraction failure that a pretrained DATE-capable model and a small transparent rule do not address. Annotate the complete spans, including modifiers such as last, and keep a held-out set. If the remaining problem is wrong year inference, DMY/MDY, event selection, or arithmetic, fine-tuning the extractor is not the first fix.

An initial decision tree:

Mostly structured dates or short phrases?
  YES -> strict known-format parser + natural-language parser.
  NO  -> do full date phrases get lost/fragmented in free text?
           YES -> try an existing DATE/temporal extractor.
           NO  -> keep extraction; diagnose the later stages.

Full correct phrase extracted, but result wrong?
  -> Fix normalization, reference time, locale or precision.
Several plausible dates, wrong event selected?
  -> Fix event linking/negation/candidate selection.
Repeated span misses on new representative examples?
  -> Compare pretrained checkpoints, then consider fine-tuning.

These are alternatives, not a to-do list.

A compact implementation contract and Python examples

Keep extracted evidence separate from accepted facts

Your example output is a good shape for a resolved event:

{"event": "relationship_start", "date": "2024-08-18"}

But an incomplete phrase should not be forced into the same finalized representation. If the input was 18 August, recording 2026-08-18 without marking the year as inferred invents a piece of user-provided evidence. Likewise, 03/04/2024 can represent different dates under different date-order conventions. A convenient intermediate candidate can preserve uncertainty:

{
  "source_text": "our anniversary was on 18 August",
  "matched_text": "18 August",
  "event_guess": "anniversary",
  "normalization": {
    "status": "needs_year",
    "month": 8,
    "day": 18,
    "year_supplied": false
  },
  "accepted_relationship_start": null
}

An intermediate record need only preserve span, source, value/unresolved reason, event and acceptance state. Offsets make clipped-span and context checks easier.

Strict known-format dates before permissive natural-language parsing

For inputs consisting exactly of ordinary Gregorian ISO calendar dates, Python has a small deterministic branch:

import re
from datetime import date

ISO_DAY = re.compile(r"\d{4}-\d{2}-\d{2}\Z")

def parse_explicit_iso(text: str):
    if not ISO_DAY.fullmatch(text):
        return None  # not this format; try another route
    try:
        return date.fromisoformat(text)
    except ValueError:
        return None  # invalid date, e.g. 2024-02-30

assert parse_explicit_iso("2024-08-18") == date(2024, 8, 18)
assert parse_explicit_iso("2024-02-30") is None

This is a parser for a candidate substring, not a recognizer that can find an ISO date anywhere in an arbitrary sentence. Extracting the candidate first is still necessary for free text. In a production interface, distinguish not_this_format from invalid_known_format: do not send a clearly malformed ISO date through an increasingly permissive series of parsers until one returns an unrelated date. The example uses None only to keep the code short.

A separate natural-language branch with a reference time

# Illustrative; install a compatible dateparser release.
from datetime import datetime
import dateparser

reference = datetime(2026, 10, 10, 12, 0, 0)

resolved = dateparser.parse(
    "two weeks ago",
    languages=["en"],
    settings={
        "RELATIVE_BASE": reference,
        "DATE_ORDER": "DMY",  # only for branches requiring this convention
    },
)
print(resolved)

Use a fixed RELATIVE_BASE in tests so output does not change tomorrow. In an application, supply the relevant reference time and timezone policy. Do not push a universal DATE_ORDER="DMY" setting into an ISO branch; that was an actual confounder in my probe. Prefer parsing known formats according to their syntax before applying locale/date-order defaults to genuinely ambiguous inputs. See the dateparser settings documentation.

A useful pipeline contract

extract(text)         -> candidate spans (text, offsets, source)
normalize(span, ctx)  -> resolved date OR explicit unresolved reason
link_event(span, ctx) -> event type / uncertainty / negation
choose(candidates)    -> accepted date OR confirm/abstain
calculate(date, now)  -> calendar-day count OR elapsed duration

Plain functions are enough initially. Recognize a few explicit event phrases; otherwise prefer needs_confirmation to inventing a date.

Normalization and candidate selection: the likely practical edge cases

Absolute, relative, incomplete, and ambiguous dates

Extraction does not supply every interpretation rule:

Phrase Necessary policy or context
2024-08-18 Strict syntax and a real calendar date
18/08/2024 Explicit date-order/locale convention
03/04/2024 Ambiguous if both MDY and DMY are allowed
last Friday Reference date/time and “last” semantics
two weeks ago Reference instant/date; whether calendar days or exact elapsed time
18 August Missing year; do not hide a guessed year
the fifth of May 2022 Support worded ordinals or return unsupported/uncertain
September 2024 Month precision, not a known exact start day

A parser can return a complete datetime even if the user supplied only month and day. That inferred year is not equivalent to an observed year. Prefer retaining explicit component/precision metadata or a separate unresolved state, rather than assuming the parser’s concrete result proves every component was supplied. In another implementation, an API exposing which date components were specified can help; chrono-node provides component information. Check behavior in the version you use.

Do not guess whether 03/04/2024 means 3 April or March 4 without a known date-order preference. Losing last changes the meaning of last Friday. Exact seconds require a reference instant and timezone, not merely a relative calendar date.

Normalizing worded ordinals can be a small rule task

Our hand-written diagnostic included the fifth of May 2022 and May fifth, 2022. The parser route used in that test failed to normalize them, even when the correct phrase was supplied. A narrow English ordinal converter, followed by calendar validation, handled them in our controlled tests. This does not establish that a new converter is needed for every application or that it generalizes to unseen prose. It illustrates a useful choice: if the extractor returns the right text, a guarded normalization rule may fix the failure more directly than additional extraction training.

Restrict ordinal conversion to recognized date patterns, and reject impossible dates such as the thirty-first of April 2024.

Multiple dates, negation, and event roles

Consider the difference between:

We started dating on 2024-08-18.
We broke up on 2024-08-18.
Our start date was 2024-08-18, not last Friday.
We booked a table for 2024-08-18.

The dates can all be extracted perfectly. Only some correspond to the event relationship_start. An initial event layer could use a few explicit, proximity-aware patterns and a safe fallback when more than one candidate competes. An event classifier may help later, but DATE labels alone are not event-role labels.

A defensible first selection policy:

  1. Retain all plausible date candidates, their original spans and nearby context.
  2. Prefer a complete, non-negated expression explicitly linked to the requested event.
  3. Treat not X, proposals, corrections, and unrelated events as distinct from an accepted start date.
  4. If competing candidates or essential missing components remain, return needs_confirmation rather than automatically selecting the first candidate.
  5. Validate the chosen calendar date before starting the duration counter.

the day before yesterday and yesterday are not interchangeable; first match, shortest match, or longest match are heuristics, not guaranteed event decisions.

What should trigger a confirmation request?

At minimum: ambiguous numeric order, missing year when a real start year is needed, unresolved relative reference, conflicting event dates, or a date that was found but not reliably connected to the event. An app does not need to interrupt the user for every minor uncertainty; explicit ISO dates can pass straight through. A targeted question such as “Is that 3 April or March 4?” is more helpful than a generic parsing error.

How to evaluate the pipeline without an expensive benchmark

Measure boundaries, not only the final counter

For each case, record the input, expected date spans/offsets, expected normalized result (or intentionally unresolved status), event role, and acceptance decision. Then compare:

  • Detection: Did a correct date expression get found?
  • Boundary: Was the whole expression found, including words such as last or before?
  • Normalization: Given the correct span and fixed context, was the resulting date right—or appropriately unresolved?
  • Event selection: Did the chosen date belong to the intended event?
  • Validation: Did invalid/ambiguous inputs trigger a safe outcome?
  • Duration: Given a correct date, did deterministic arithmetic match the chosen convention?

Do not confuse a span overlap metric with exact match. An extracted Friday that overlaps the gold last Friday is a hit under some overlap scores but loses meaning. Similarly, a first-candidate strategy may score well on positive single-date phrases and fail badly on negative or competing-date inputs.

Two cheap controls that locate the real problem

Gold-span (oracle) control: Hand the normalizer the exact correct date phrase. If the calculator remains wrong, fixing the extractor cannot by itself repair the final result.

Gold-date control: Hand the event-selection/duration stage a verified, normalized date. If it still chooses the wrong event or count, inspect event logic or time arithmetic rather than models or parsing.

These controls are usually more informative than immediately comparing ten more checkpoints. You can implement them by passing a hand-written candidate or resolved value into the same functions; no training is needed.

Suggested tiny regression set

Include your five originals and a few changes that each isolate a distinct failure: Sept. 18, 2024; last Friday versus Friday; the day before yesterday; an invalid leap-day or April 31; 03/04/2024 with unspecified date order; 18 August without a year; the negated two-date sentence; a different event; and an input with no temporal phrase. Record whether the expected outcome is a specific day, an alternative set, or abstention.

For future fine-tuning, reserve unseen phrasing/sources, separate rule-development cases, and annotate language, date order and reference time. Evaluate exact spans, downstream decisions, cost and latency.

Independent experiments: observations, controls, and limitations

I ran small diagnostics because I wanted to check where the design separation has practical consequences. These were not your code and do not support model rankings; several rules were developed on the same hand-written cases.

CPU parser comparison

A Colab CPU run compared dateparser 1.2.2 and 1.4.3 on 20 selected cases, after an earlier Windows/Python 3.9 test of 1.2.2. It completed without test-time exceptions. In this setup, parser long-text search sometimes returned only Friday for last Friday, silently supplied a year for 18 August, missed a date in a sentence containing both 2024-08-18 and not last Friday, or returned We as an additional candidate. The newer split and ngram paths did not solve all these chosen cases. That does not imply every version or configuration has these behaviors for typical usage; it explains why a baseline should be inspected case by case.

Small GPU extraction/pipeline comparison

A separate Colab GPU run used 30 chosen English cases, 31 manually labeled date spans, and 10 cases with a labeled relationship start. The run completed with no inference errors using pretrained GLiNER, a specialist temporal BERT via a generic token-classification pipeline, dateparser search, hand-built regexes, and a simple hybrid. The GLiNER inference was zero-shot with respect to these user-labeled date types; it was not custom fine-tuned for the dataset. The downstream parser and event rules were shared across methods.

This table counts overlapping gold spans, not exact-span accuracy or unbiased model performance:

Extractor path Gold spans overlapped /31 Original correct start selections /10
dateparser long-text search 19 4
Hand-built regex + parser 27 3
GLiNER (date) + parser 27 5
GLiNER (date/time/duration) + parser 28 5
Temporal BERT generic pipeline + parser 20 5
Simple hybrid + parser 30 4
Correct spans supplied directly (oracle) 31 5

The regex was authored with knowledge of the examples. Partial spans count as overlap. The BERT experiment omitted the author’s recommended alignment/voting postprocessing and therefore does not evaluate the complete temporal-tagger method. The hybrid had simple, untuned candidate priority. Do not rank the models by these rows.

What changed when I fixed normalization only?

The shared normalizer had used DATE_ORDER="DMY" for all candidates, including ISO strings. That produced wrong interpretations in some cases—even with the correct spans supplied by an oracle. I kept the saved extractor outputs fixed and replayed the downstream normalization offline: first route complete YYYY-MM-DD candidates through strict ISO parsing, then add a narrow validated rule for English worded ordinals.

Same frozen candidate spans Original GPU result Strict ISO replay ISO + ordinal replay
GLiNER (date) 5/10 9/10 10/10
GLiNER (date/time/duration) 5/10 9/10 10/10
Simple hybrid 4/10 8/10 9/10
Oracle spans 5/10 9/10 10/10

These are not new model predictions. They are conditional, post-hoc replays on ten hand-selected relationship-start cases. They show that downstream code can dominate the final outcome even when an extractor already has the necessary span. 10/10 does not mean GLiNER or the end-to-end system has 100% accuracy on unseen data. The remaining hybrid case illustrates candidate precedence (the day before yesterday versus yesterday), not necessarily a missing model capability.

An ordinal-date CPU check

A subsequent CPU probe compared existing parsing paths with a guarded English ordinal-normalization rule on 28 hand-designed positive/negative/control strings. The tailored rule met the specified expectations for all 28 examples, including valid leap-day phrasing and rejection of invalid dates. This is a development-panel check, not held-out generalization.

The result supports measuring extraction, normalization, and event selection independently before fine-tuning.

Date/time arithmetic and the next extension points

Your decision to leave the arithmetic to deterministic code is especially helpful because a “duration” can mean two different things:

  • Calendar days together: subtract local calendar dates. If the displayed counter says “days since we started,” define whether the start date is day 0 or day 1.
  • Exact elapsed days/hours/minutes/seconds: subtract two timezone-aware instants. If the user entered only 2024-08-18, there is no known start hour, minute, second, or timezone; choosing midnight is a product convention, not information extracted by the model.

For example, a minimal date-only calculation is:

from datetime import date
start = date(2024, 8, 18)
today = date(2026, 10, 10)   # test date only
calendar_days = (today - start).days
print(calendar_days)

For an exact seconds counter, the input must include or explicitly assume a start time and timezone. Daylight saving transitions can make elapsed hours differ from calendar_days * 24. “Months together” and “years together” also require conventions for varying month lengths, leap-day anniversaries, and rounding; they are not reliably obtained by dividing seconds by a constant.

Later extensions—ranges, recurring events, multiple languages and calendars—need their own representations and tests.

For privacy, log only necessary test IDs/spans; use synthetic or consented/anonymized data for public benchmarks.

Useful references and an optional roadmap

Practical learning and implementation

Adjacent discussion and deeper context (only if you need it)

A low-effort sequence could be: (1) compute durations from verified ISO dates; (2) normalize short natural-language expressions with a reference date; (3) compare the exact spans returned by a pretrained DATE extractor and a parser search; (4) handle one competing-date/negation example; (5) collect representative new failures before considering fine-tuning. You can stop at any step if that already meets the intended scope. The HF extraction exercise can be done first if learning is the priority.

Returning to your question

Yes: DATE-capable NER/token classification is a sound first Hugging Face exercise here. Keep the separate parser and deterministic calculator. Fine-tune a temporal extractor only when you can point to recurring span extraction errors on representative examples. If the extracted phrase is already correct, check normalization and event selection first. A small baseline plus a few contrast cases will tell you where extra model capacity actually helps—without changing the central idea of your project.