Imtiaz Mashrafee

AI Engineer at Data Solution-360
Computer Science, BRAC University
Dhaka, Bangladesh

Imtiaz MashrafeeAI Engineer

I build AI systems and test what they do when inputs go wrong.

One of five cases. Choose one and the check below shows what the system does with that input.

  1. Input“I need an email about my job assessment.”AI email writer
  2. CheckToo vague to write fromcaught
  3. ControlAsk a question first
  4. OutcomeNo email is writtenUntil the request is clear
  1. Input2 seats requested, 3 of 3 takenRide pooling
  2. Check3 + 2 is more than 3caught
  3. ControlRefuse the request
  4. OutcomeNothing changesStill 3 of 3 seats
  1. InputCell Count 1.2 x 1043Extraction
  2. CheckMatches no known number formcaught
  3. ControlLeave the row out
  4. OutcomeNot guessed, not keptThe rest of the report goes through
  1. InputDraft answer scored 0.67Retrieval
  2. Check0.67 is below the 0.72 thresholdcaught
  3. ControlOne repair attempt
  4. OutcomeChecked againBefore the answer is used
  1. Input1 seat requested, 0 of 3 takenRide pooling
  2. Check0 + 1 is within 3passes
  3. ControlNo intervention needed
  4. OutcomeAccepted1 of 3 seats taken
Each request is checked before the system acts. These cases come from the four projects below.

System behavior bench

Choose which system to operate. The page moves to its chapter. Arrow keys move between systems.
Ride pooling capacity guard
Recorded project run

Recorded from the project's own test run at the commits before and after the guard, with an in-memory database mock.

0123456 A B C 2 seats send capacity 3 committed 3 of 3409refused3 + 2 > 3
0123456 A B C 2 seats send capacity 3 committed 3 of 3409refused3 + 2 > 3
Retrieval verify and repair
Recorded project run

Final scores recorded in the committed evaluation runs (mock generator, not answer quality).

Recorded scores (drafts from a stand-in generator)
0.1.2.3.4.5.6.7.8.91.0 required quality 0.72 recorded run, query q3: Compare CNNs and Transformers 0.67still below the requirement0.67 < 0.72, after one repairstepsfind sourcesrank themtrimset contextwrite draftcheck qualityone retryone retry: more weight on exact keyword matches, stricter trimming,a tighter answer template and less randomness, then check again
0.1.2.3.4.5.6.7.8.91.0 required quality 0.72 recorded run, query q3: Compare CNNs and Transformers 0.67still below the requirement0.67 < 0.72, after one repairstepsfind sourcesrank themtrimset contextwrite draftcheck qualityone retryone retry: more weight on exact keywords,stricter trimming, a tighter template
Extraction report parsing
Recorded project run

Observed OCR misreads recorded in the project README (self-reported evaluation); raw OCR outputs are not committed.

OCR line
testunitrangeflag
testvalueunitrangeflagstatus
Hemoglobin10.5g/dL12.0 - 16.0Lkept
Platelet Count1250010^3/uL10,000 - 15,000N/Akept
omitted: the value did not parseomitted
Glucose180mg/dL70 - 100Hkept
Albumin3.7g/dL3.5 - 5.0N/Akept
4 of 5
rows kept
value did not match any form
Ride pooling Blocks requests over capacity, except two at once.
  1. Start: a car with 3 seats
    1 of 3 taken
  2. A request for 2 seats
    3 of 3: accepted
  3. 2 more seats requested
    would be 5 of 3: refused car stays 3 of 3 · HTTP 409
  4. Back at 1 of 3: two requests for 2 seats at the same moment
    request 1 reads “1 of 3”: fits, so it books
    request 2 reads “1 of 3”: fits, so it books
    5 of 3 observed
Engineering mechanism

Every status change goes through one transition table (REQUESTED, MATCHED, DRIVER_ARRIVED, STARTED, COMPLETED). The capacity guard runs inside the accept transaction and adds up the seats of every active pool of the vehicle; a refusal is HTTP 409. Two simultaneous accepts can both pass because the check reads, then writes (seen under an in-memory database mock).

Another request for 2 seats

ABCcapacity 3
0123456

The guard adds up every active pool of the vehicle: 3 committed + 2 requested > 3.

Refused3 + 2 is more than 3, so the car stays at 3 of 3. HTTP 409

Where the check sits in a request's life
  1. REQUESTED
  2. MATCHEDaccept: capacity guard
  3. DRIVER_ARRIVED
  4. STARTED
  5. COMPLETED

Every change goes through one transition table. Cancellation is allowed before the ride starts.

from project docsThe check runs when a driver accepts, not when the rider asks.

Count the whole car, not each group

  1. each group on its ownlooks finepasses
  2. the whole car added up3 + 2 is more than 3refused
recorded runA check on each group passed while the car was over capacity, so the check moved up to the whole vehicle.

Before and after the seat check was added

ABCcapacity 3
0123456

Without the guard: 5 of 3.

ABcapacity 3
0123456

With the guard: 3 of 3.

recorded runRecorded before and after the change. Mock database, not a real one.

Two requests at the same moment

Each request checks the seats, sees room, then books. Neither sees the other, so both get in.

accept 1reads 1 of 3passeswrites

accept 2reads 1 of 3passeswrites

ABCcapacity 3
0123456

Both passed the check: 5 of 3.

recorded runRequests accepted at the same moment are not covered by the check. Seen under the mock; a real database was not run.
Email Checks the request before the AI writes.
“I need an email about my job assessment.”
1 fact given · tone: professional
✓Basicsby code: purpose, fact, tone
✕Enough to write from?judged by a model
Ask a question first the AI stops before it writes
Write the email
without the checkwrote an email anyway · 2.33 of 5, model-scored
with the checkasked first · 5 of 5, model-scored
Engineering mechanism

Request validation in code (intent not empty, at least one non-empty fact, tone from a fixed list), then the input checker, a model call that decides whether there is enough usable information (not run in this page), then a clarify or generate branch. An evaluation runner scores both strategies on ten recorded scenarios. This is scenario 3; the scores are recorded judge outputs, and why the checker decided as it did was not saved.

First check: does the request have the basics?

  • it says what the email is for
  • it gives at least one fact no facts given
  • the tone is one of professional, polite, formal, friendly, empathetic, urgent but respectful, casual but respectful

Shown: a request with every fact removed. It stops here, before the AI is asked to write anything.

Technical message

List should have at least 1 item after validation, not 0

from project docsThe same rules run on the stage above when you edit a request.

Is there enough to write from?

a request with the basics

decides if there is enough to write frominput checker: a model call, not run in this page

asks a question firstrecorded for scenario 3: writing stops until the request is clearer

writes the emailthe other branch

recorded runThe decision is the model's. The code decides only what reaches it.

Tested on 10 recorded scenarios

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10

9 of 10: asked or wrote, as expected. Average score 4.7 out of 5, against 4.2 for a single prompt. One run, scored by a model.

recorded runOne saved run of ten scenarios, not a live evaluation.

One unclear request still got through

request
Request a deadline extension for my final project.
should have
asked first
actually
wrote the email
recorded runThe checker is a model: in scenario 4 it did not ask when it should have.
Retrieval Checks the first answer before returning it.
“Compare CNNs and Transformers”
Passages are found, then the AI writes a first answer. Quality is a rule-based score from 0 to 1.
First draftNot good enough yet
One more attempt, once: it searches again, leaning on exact keywords, and writes more strictly
After the retryStill belowreturned marked not accepted
A retry is not a guarantee: 2 of 4 recorded questions passed first time.
Engineering mechanism

The route is retrieve, rerank, prune, contextualize, generate, verify, and repair below the threshold. The verify score (0 to 1) is rule-based: length 20%, relevance 30% (how many of the question's words the answer uses), coherence 25%, completeness 15% and repetition 10%. It does not compare the answer with the retrieved passages. The default accept threshold is 0.72. The first draft's score was not recorded; q3's final score after repair was 0.67. Drafts came from a mock generator.

Is the draft good enough?

one more attemptgood enoughrequired 0.72draft 0.67
0.51.0

Below the required qualityThe draft scored 0.67; the requirement is 0.72.

runs in browserThe acceptance rule on the stage is this comparison, run by the same code.

A second attempt, once

  1. first draft
  2. below the requirement
  3. one repair attempt
  4. checked again

the system does not just return its first answer

Engineering route
  1. retrieve
  2. rerank
  3. prune
  4. contextualize
  5. generate
  6. verify

Below the threshold: one repair attempt, then verify again.

Repair attempt: bm25 0.5 to 0.7, vector 0.5 to 0.3, prune overlap 1 to 2, constrained template, temperature 0.2 to 0.1.

from project docsFinding sources, pruning and drafting come before the check; the repair attempt comes after it.

What happened to four recorded questions

q10.82accepted
q50.72accepted
q30.67tried again, still below
q40.655tried again, still below
recorded runThe path each question took is saved with the evaluation runs.

A repair is an attempt, not a guarantee

required 0.72q3 0.67q4 0.655
0.51.0
recorded runBoth repaired answers stayed under the requirement. A mock wrote the drafts, so the scores describe the checking, not answer quality.
Extraction Turns messy readings into structured fields.
A line read from a photographed lab report
RP<8.5mg/dL<1.0N/A
testvalueunitrangeflag
✓looks like a value: kept
the report says<0.5
it was read as<8.5
Looks valid. Nothing flags it.
Engineering mechanism

The route is OCR, a lab-report check, metadata, the value parser and row omission. The parser accepts a plain number, a qualified number (such as <8.5) or a scientific value, and leaves out any row whose value matches no known form, keeping the raw line for rows that parse. A misread digit that still forms a valid number passes, and there is no confidence measure. The README records <8.5 read in place of <0.5.

How a value is read

test
CRP
value
<0.5
unit
mg/dL
range
<1.0
flag
N/A

The value is accepted only if it has a shape the parser knows: a plain number, a qualified number or a scientific value. Nothing else.

from project docsA cautious parser: it accepts what it knows and refuses the rest.

A reading it cannot make sense of

test
Cell Count
value
1.2 x 1043
unit
10^3/µL
range
1.0 - 2.0
flag
N/A

row left outIt is not guessed and not kept.

test dataThe README records this misread, 1.2 x 10^3 read as 1.2 x 1043.

One misread is caught, one is not

  1. 1.2 x 1043does not look like a numberrow left out
  2. <8.5looks like a valid numberrow kept
test dataThe first is caught. The second is not: that is the next step.

It looks valid, but it is wrong

test
RP
value
<8.5
unit
mg/dL
range
<1.0
flag
N/A

parsedlooks like a valid value

should read<0.5

was read as<8.5

keptNothing flags it, so a reader would never know.

test dataThere is no confidence measure, so a valid-looking wrong value passes.

Other systemsselect one to switch

Ride-pooling lifecycle with guarded transitionsoverbooking

How do you stop a car being overbooked?

A ride-sharing backend that refuses a booking when the car would end up with more riders than seats. It does not yet hold when two bookings arrive at the same instant.

My part: the ride service, the request steps and the seat check. Sole contributor in the commit history (all 125 commits).

What goes wrong: Ride-pooling lifecycle with guarded transitions

Goes wrong
Several riders can ask for seats in the same vehicle, and each request looks fine on its own.
Why it matters
If every request is accepted separately, the car can end up with more riders than it has seats.
What it does
It adds up the seats already taken across the whole vehicle and refuses a request that would not fit.
In engineering terms

A shared-ride service where a request moves only along a transition table, with a capacity check across a vehicle's pools that holds for sequential accepts; concurrent accepts remain a known gap.

Pooling riders into shared trips under constraints while several actors change state at once, with an auditable history.

TypeScript, Next.js, Express, PostgreSQL

How it responds: Ride-pooling lifecycle with guarded transitions

Select a step to see what it does and why it exists.

  1. A rider asks for seats Request

    A passenger creates a request with a seat count, and the server computes an integer-paisa fare.

    Decision. Integer paisa for money.

  2. Every request follows fixed steps Transition table

    Every status change goes through one table: REQUESTED, MATCHED, DRIVER_ARRIVED, STARTED, COMPLETED, with cancellation allowed before the ride starts.

    Decision. One transition table for every status change.

    more about Transition table
    Why it exists
    So no route can skip a state.
    Evidence

    Verified 71 written test cases exist in the repository. They were counted, not run for this page.

    Source
    apps/api/src/common/rideLifecycle.ts
  3. A driver accepts the request Conditional accept

    A driver accepts a request. A conditional update succeeds only if the request is still REQUESTED.

    Decision. Conditional (compare-and-set) updates for single-row races.

    more about Conditional accept
    Limitation
    It covers a single row, not the vehicle's total across pools.
  4. Riders are grouped into one trip Pool matching

    The service joins an existing pool when routes are compatible, or founds a new pool.

    more about Pool matching
    Why it exists
    Pools form by route compatibility, within pickup and destination distance limits.
  5. Blocks overbooking Capacity guard

    Adds up the seats committed across every active pool of the vehicle (open or locked) and compares that total plus the new request with the vehicle's capacity.

    Decision. Guard the vehicle's total, not each pool.

    more about Capacity guard
    Why it exists
    A per-pool check passed while the vehicle as a whole was over capacity, so the guard was moved to the vehicle level.
    In and out
    Requested seats (1 to 6) and the vehicle's active pools. Accept, or HTTP 409 and a conflict status event.
    Failure
    Two accepts at the same moment can both read a total under capacity before either writes.
    Control
    The accept transaction refuses when the total plus the request exceeds capacity.
    Evidence

    Verified Recorded before and after the guard commit: 5 of 3 seats without it, 3 of 3 with it. In-memory database mock, not PostgreSQL.

    Partially verified Two accepts at the same moment both passed the guard under the mock, reaching 5 of 3. PostgreSQL behavior was not run.

    Limitation
    Concurrent accepts across pools are not covered by this guard.
    Source
    apps/api/src/modules/rides/rides.service.ts, accept transaction
  6. Every change is logged Status events

    Every transition leaves a status event, and a conflict outcome is written as an event after the rollback.

    Decision. Success events committed with state changes, and conflict events written after a rollback.

Decision: Ride-pooling lifecycle with guarded transitions

Count seats for the whole car, not for each group of riders, at the moment a driver accepts a request.

In engineering terms

Guard the vehicle's total seats across pools, not each pool, inside the accept transaction.

What happened: Ride-pooling lifecycle with guarded transitions

Recorded test run: without the check a 3-seat car held 5 riders; with it, 3 of 3. This ran against a stand-in (mock) database, not a real one.

Evidence detail Checked by me

71 written test cases exist in the repository. They were counted, not run for this page.

Where it still fails: Ride-pooling lifecycle with guarded transitions

Two requests accepted at the same moment can both pass the check, which would leave the car at 5 of 3. Seen only against a stand-in (mock) database; a real database was not run.

In engineering terms

Concurrent accepts across pools are not covered by the guard. Observed under an in-memory mock; PostgreSQL was not run.

Next systemWhen should an AI not start writing?

Two-stage email drafting pipelineunclear requests

When should an AI not start writing?

An AI email writer that first checks whether the request makes sense, and asks a question instead of writing when it does not.

My part: the design of both approaches, the request checker, the scoring framework, the ten test scenarios and the demo app. Sole contributor in the commit history.

What goes wrong: Two-stage email drafting pipeline

Goes wrong
A single prompt writes a confident email even when the request is vague, contradictory or missing the facts it needs.
Why it matters
A confident email looks finished, so nobody notices when it says the wrong thing.
What it does
It checks the request first, and asks a question instead of writing when there is not enough to go on.
In engineering terms

Compares a single prompt with a pipeline that checks inputs before it generates, and scores both against a fixed set of scenarios.

A single prompt writes a confident email even when the request is unclear, contradictory, padded with irrelevant facts or contains instruction-like text.

Python, Streamlit, LLM provider APIs, Pydantic

How it responds: Two-stage email drafting pipeline

Select a step to see what it does and why it exists.

  1. Checks the request has the basics Request validation

    Checks the request in code before any model is called: the intent is not empty, there is at least one non-empty fact, and the tone is one of the allowed values.

    Decision. Validate in code first.

    more about Request validation
    Why it exists
    Cheap, deterministic failures should never reach a model.
    In and out
    Intent, key facts and tone. A valid request, or an error that stops the run.
    Failure
    Code can only check shape. It cannot tell whether the request is clear enough.
    Control
    A validation error is raised before the input checker runs.
    Evidence

    Verified The messages and rules shown here match the original Python on a grid of test inputs. Checked with the pydantic version recorded in the golden file.

    Limitation
    Whether a request is clear is decided by a model, not by this stage.
    Source
    src/schemas.py, model_b_approach/pipeline.py
  2. Decides if there is enough to write from Input checker

    A model call that decides whether the request has enough usable information. It returns a status, a canonical intent, the usable facts and, when needed, a clarifying question.

    Decision. Make clarification a normal outcome of the pipeline.

    more about Input checker
    Why it exists
    Ask first instead of guessing.
    In and out
    A valid request. Ready with a canonical intent and usable facts, or clarify with a question.
    Failure
    It is a model, so it can miss: in one recorded scenario a missing extension length did not trigger a question.
    Control
    A clarify status returns the question and no email is generated.
    Evidence

    Partially verified Ten scenarios, one run: the gated pipeline asked first in 2 of the 3 scenarios that expected a question. Judge and generator models were not recorded; the result was not reproduced.

    Limitation
    Not run in this page. Outcomes are recorded, and any edit retires them.
    Source
    model_b_approach/input_checker.py
  3. Asks a question, or writes Clarify or generate

    On a clarify status the pipeline returns the question and no email. Otherwise the generator writes the email.

    Decision. Make clarification a normal outcome of the pipeline.

    more about Clarify or generate
    Source
    model_b_approach/pipeline.py
  4. Scores the results on test scenarios Evaluation runner

    A separate runner sends scenarios through both strategies and scores the results with three judge metrics and structural checks. It is not part of the pipeline.

    more about Evaluation runner
    Evidence

    Partially verified The committed summary averages are 4.7 for the gated pipeline and 4.2 for the baseline, on a scale of 1 to 5. Ten scenarios, one run; the judge and generator models are not recorded.

    Source
    evaluate.py

Decision: Two-stage email drafting pipeline

Check the request with ordinary code first, then let a model judge whether there is enough to write from.

In engineering terms

Validate in code first, then let a model decide whether to ask a question before any email is written.

What happened: Two-stage email drafting pipeline

Tested on 10 recorded scenarios: the pipeline averaged 4.7 out of 5 and a single prompt 4.2. One run, scored by a model that was not recorded.

Evidence detail Checked by me

Ten test scenarios are saved in the repository, with the recorded outputs and the scores a judge model gave them.

Where it still fails: Two-stage email drafting pipeline

One unclear request still got through. For "Request a deadline extension for my final project" it wrote an email when it should have asked.

In engineering terms

Ten scenarios and one run. The checker is a model: in one recorded scenario it did not ask when it should have.

Next systemWhen should an AI trust its first answer?

Retrieval over research papers with verify and repairweak answers

When should an AI trust its first answer?

An AI that answers questions about research papers, scores its own draft for quality, and gets one retry when the draft falls short.

My part: the whole system: reading the papers, finding sources, writing and checking answers, and the code that tests it. Sole contributor in the commit history (all six commits).

What goes wrong: Retrieval over research papers with verify and repair

Goes wrong
A first draft can be too short, drift away from the question, read badly or stop mid-thought, and nothing would tell the reader.
Why it matters
An answer returned without any check is only as good as the first attempt.
What it does
It scores each draft with simple quality rules. Below the required quality, it makes one more attempt with tighter settings and scores again.
In engineering terms

A plan-driven pipeline that retrieves, scores the draft answer with a rule-based quality check and makes one repair attempt when the score is below a threshold.

Answering questions over research papers, including tables, with answers that are checked rather than trusted.

Python, FastAPI, DuckDB, sentence-transformers

How it responds: Retrieval over research papers with verify and repair

Select a step to see what it does and why it exists.

  1. Works out what the question needs Query plan

    A plan is chosen from the question's intent and then executed.

    Decision. A plan keyed on query intent instead of one fixed chain of steps.

  2. Finds supporting passages Hybrid retrieval

    Two indexes, BM25 and an in-memory embedding store using cosine similarity, behind an ensemble retriever.

    Decision. Hybrid lexical and vector retrieval.

  3. Drops sentences that do not help Sentence-level prune

    Irrelevant sentences are pruned before the context is built.

    Decision. Sentence-level pruning before contextualization, so the generator sees less irrelevant text.

  4. Checks the answer is good enough Acceptance rule

    Compares the verify score of a draft answer with the acceptance threshold. At or above the threshold the answer is accepted and repair is skipped.

    Decision. Make acceptance a threshold that a request can override.

    more about Acceptance rule
    Why it exists
    Acceptance is an explicit rule with a number, not an assumption.
    In and out
    A verify score from 0 to 1 and the threshold (0.72 by default). Accepted, or one repair attempt.
    Failure
    The verify score is only as meaningful as the verifier and the generator; the committed runs use a mock generator.
    Control
    A score at or above the threshold is accepted; below it, one repair runs.
    Evidence

    Verified Committed runs record the final score, whether repair ran and whether the answer was accepted. Mock generator; scores describe harness behavior, not answer quality.

    Limitation
    No claim is made about answer quality or real model behavior.
    Source
    rag_papers/retrieval/router_dag.py, run_plan and exec_repair
  5. One more attempt if it falls short Repair step

    Makes one more pass with tighter settings: more weight on BM25 (0.5 to 0.7) and less on vectors (0.5 to 0.3), a stricter prune, a constrained template and a lower temperature (0.2 to 0.1), then verifies again.

    Decision. One repair only.

    more about Repair step
    Why it exists
    A bounded second attempt has a fixed cost, and acceptance is decided by the same rule.
    In and out
    A draft below the threshold. A new draft and a new verify score.
    Failure
    After the single attempt the answer can still be below the threshold and is then returned as not accepted.
    Control
    The maximum number of repairs is 1.
    Evidence

    Partially verified In the committed runs that produced scores, every query that used repair ended below 0.72 and was not accepted. Mock generator; scores describe harness behavior.

    Limitation
    Whether a repair improves an answer is not claimed.
    Source
    rag_papers/retrieval/router_dag.py, exec_repair

Decision: Retrieval over research papers with verify and repair

Do not return the first draft. Score it, accept it at the required quality, and allow one bounded repair attempt below it.

In engineering terms

Score each draft with the rule-based verifier, accept at a threshold, and make one bounded repair attempt below it.

What happened: Retrieval over research papers with verify and repair

Four recorded queries: two accepted on the first draft, two tried again and still below the requirement. The drafts came from a stand-in (mock) generator.

Evidence detail Checked by me

Saved evaluation runs record which path each question took through the pipeline.

Where it still fails: Retrieval over research papers with verify and repair

Repair is an attempt, not a guarantee: both repaired answers stayed below the requirement. The drafts came from a stand-in (mock) generator and the score is a simple rule-based quality check, so it shows how the checking behaves, not how good or correct the answers are.

In engineering terms

The committed runs use a mock generator, so scores describe harness behavior and not answer quality. Repair is not a guarantee.

Next systemWhat if a reading looks valid but is not?

Speech and report extraction that does not invent valuesmisread values

What if a reading looks valid but is not?

A system that transcribes speech and pulls structured data out of photographed lab reports, leaving out any value it cannot read instead of guessing.

My part: the whole service: its web endpoints, the code that reads and sorts report text, and the tests and sample data. Sole contributor in the commit history (all 51 commits).

What goes wrong: Speech and report extraction that does not invent values

Goes wrong
Reading text from a photographed lab report sometimes misreads a value, and a misread can still look like a perfectly good number.
Why it matters
A wrong lab value that looks valid is worse than a missing one, because nothing warns the reader.
What it does
It turns report text into structured fields and leaves out values it cannot read. The page also shows the case it cannot catch.
In engineering terms

Transcribes speech and extracts structured rows from photographed lab reports, omitting rows whose values cannot be parsed and returning zero rows for documents that are not lab reports.

Turning messy audio and photographed documents into structured data without making values up.

Python, FastAPI, faster-whisper, Tesseract

How it responds: Speech and report extraction that does not invent values

Select a step to see what it does and why it exists.

  1. Reads the text from the photo OCR

    A provider reads a photographed report into text lines. A mock provider is the default and Tesseract is the real one.

    Decision. Provider adapters with mock defaults.

    more about OCR
    Limitation
    OCR degrades on rotated or angled images, and raw outputs are not committed.
  2. Checks it is a lab report Lab report check

    A conservative detector decides whether the text looks like a lab report. If not, the type is unknown and zero rows are returned.

    Decision. Conservative non-lab detection.

    more about Lab report check
    Evidence

    Partially verified A receipt returned the type unknown and 0 rows. Tested on one image; self-reported.

  3. Picks out patient, date and lab Metadata

    Reads patient, date and lab fields where present. With the header cropped out, metadata is null.

    more about Metadata
    Evidence

    Partially verified With the header cropped out, metadata is null and 4 rows are recovered. Self-reported README tables.

  4. Turns report text into fields Value parser

    Reads the value column as a plain number, a qualified number, a range or a scientific value. Anything that matches none of those forms returns no value.

    Decision. Reject malformed values, and do not correct OCR using knowledge of the source document.

    more about Value parser
    Why it exists
    A value that does not match a known form is not repaired or guessed.
    In and out
    The value text of one row. A value with its kind, operator, numbers and the raw text, or none.
    Failure
    A wrong value that is still well formed parses as valid, for example <8.5 read in place of <0.5.
    Control
    No match returns none, and the row is omitted.
    Evidence

    Partially verified The README records Tesseract reading CRP <0.5 as RP <8.5 and 1.2 x 10^3 as 1.2 x 1043. Self-reported evaluation; raw OCR outputs are not committed.

    Limitation
    There is no confidence measure, so the parser cannot detect a misread that is still a valid number.
    Source
    app/services/value_normalizer.py
  5. Leaves out values it cannot read Row omission

    Drops a row whose value does not parse. A row that does parse keeps its raw line.

    Decision. Omit rather than guess.

    more about Row omission
    Why it exists
    A made-up value is worse than a missing row.
    In and out
    A parsed value, or none. A result row with its raw line, or no row.
    Failure
    An omitted row is silent: the response does not list what was dropped.
    Control
    A value of none returns no row.
    Evidence

    Partially verified Recorded examples: with the header cropped out, 4 rows are recovered; a receipt returns the type unknown and 0 rows. Self-reported README tables.

    Limitation
    The OCR text of dropped rows is not returned.
    Source
    app/services/report_parser.py

Decision: Speech and report extraction that does not invent values

Accept only value shapes it knows. Leave out any row that does not match, and keep the raw line for rows that do.

In engineering terms

Reject a value that does not match a known form, omit that row, and keep the raw line on rows that parse.

What happened: Speech and report extraction that does not invent values

Two recorded misreads: one is caught and its row left out, the other looks valid and is kept. Test recordings and synthetic report images are saved in the repository.

Evidence detail Checked by me

Test recordings and synthetic report images are saved in the repository.

Where it still fails: Speech and report extraction that does not invent values

It cannot tell a wrong value that looks valid. A reading of <8.5 in place of <0.5 is parsed and kept, and nothing flags it. There is no confidence measure.

In engineering terms

A misread that is still a valid number parses as valid. There is no confidence measure.

All work

What you tested

Nothing tested yet. Operate any system above and this page writes down what you tried, what held and what did not.