
AI Engineer at Data Solution-360
Computer Science, BRAC University
Dhaka, Bangladesh
Imtiaz MashrafeeAI Engineer
I build AI systems and test what they do when inputs go wrong.
- Input“I need an email about my job assessment.”AI email writer
- CheckToo vague to write fromcaught
- ControlAsk a question first
- OutcomeNo email is writtenUntil the request is clear
- Input2 seats requested, 3 of 3 takenRide pooling
- Check3 + 2 is more than 3caught
- ControlRefuse the request
- OutcomeNothing changesStill 3 of 3 seats
- InputCell Count 1.2 x 1043Extraction
- CheckMatches no known number formcaught
- ControlLeave the row out
- OutcomeNot guessed, not keptThe rest of the report goes through
- InputDraft answer scored 0.67Retrieval
- Check0.67 is below the 0.72 thresholdcaught
- ControlOne repair attempt
- OutcomeChecked againBefore the answer is used
- Input1 seat requested, 0 of 3 takenRide pooling
- Check0 + 1 is within 3passes
- ControlNo intervention needed
- OutcomeAccepted1 of 3 seats taken
System behavior bench
Recorded project runRuns in browser
Recorded from the project's own test run at the commits before and after the guard, with an in-memory database mock.
A port of the capacity guard arithmetic running in this page. Checked against the recorded harness states.
Recorded project runRuns in browser
Outcomes recorded in the project's committed evaluation: ten synthetic scenarios, one run; the judge and generator models were not recorded.
A port of the project's request validation running in this page, checked against the original Python. The input checker is a model call and is not run here.
Recorded project runRuns in browser
Final scores recorded in the committed evaluation runs (mock generator, not answer quality).
The acceptance rule and repair routing from the repository's router, running in this page. This models control flow only: no retrieval or generation happens.
your assumption: the score after repair is not recorded
Recorded project runRuns in browser
Observed OCR misreads recorded in the project README (self-reported evaluation); raw OCR outputs are not committed.
A port of the repository's row and value parsing running in this page, checked against lines run through the original Python (see the golden tests).
Capacity guard
- What it does
- Adds up the seats committed across every active pool of the vehicle (open or locked) and compares that total plus the new request with the vehicle's capacity.
- Why it exists
- A per-pool check passed while the vehicle as a whole was over capacity, so the guard was moved to the vehicle level.
- In and out
- Requested seats (1 to 6) and the vehicle's active pools. Accept, or HTTP 409 and a conflict status event.
- Decision
- Guard the vehicle's total, not each pool.
- What it does not cover
- Two accepts at the same moment can both read a total under capacity before either writes.
- Evidence
VerifiedRecorded before and after the guard commit: 5 of 3 seats without it, 3 of 3 with it. In-memory database mock, not PostgreSQL.
Partially verifiedTwo accepts at the same moment both passed the guard under the mock, reaching 5 of 3. PostgreSQL behavior was not run.
- Limitation
- Concurrent accepts across pools are not covered by this guard.
- Source
- apps/api/src/modules/rides/rides.service.ts, accept transaction
Value parser
- What it does
- Reads the value column as a plain number, a qualified number, a range or a scientific value. Anything that matches none of those forms returns no value.
- Why it exists
- A value that does not match a known form is not repaired or guessed.
- In and out
- The value text of one row. A value with its kind, operator, numbers and the raw text, or none.
- Decision
- Reject malformed values, and do not correct OCR using knowledge of the source document.
- What it does not cover
- A wrong value that is still well formed parses as valid, for example <8.5 read in place of <0.5.
- Evidence
Partially verifiedThe README records Tesseract reading CRP <0.5 as RP <8.5 and 1.2 x 10^3 as 1.2 x 1043. Self-reported evaluation; raw OCR outputs are not committed.
- Limitation
- There is no confidence measure, so the parser cannot detect a misread that is still a valid number.
- Source
- app/services/value_normalizer.py
- What it does
- Drops a row whose value does not parse. A row that does parse keeps its raw line.
- Why it exists
- A made-up value is worse than a missing row.
- In and out
- A parsed value, or none. A result row with its raw line, or no row.
- Decision
- Omit rather than guess.
- What it does not cover
- An omitted row is silent: the response does not list what was dropped.
- Evidence
Partially verifiedRecorded examples: with the header cropped out, 4 rows are recovered; a receipt returns the type unknown and 0 rows. Self-reported README tables.
- Limitation
- The OCR text of dropped rows is not returned.
- Source
- app/services/report_parser.py
Input checker
- What it does
- Checks the request in code before any model is called: the intent is not empty, there is at least one non-empty fact, and the tone is one of the allowed values.
- Why it exists
- Cheap, deterministic failures should never reach a model.
- In and out
- Intent, key facts and tone. A valid request, or an error that stops the run.
- Decision
- Validate in code first.
- What it does not cover
- Code can only check shape. It cannot tell whether the request is clear enough.
- Evidence
VerifiedThe messages and rules shown here match the original Python on a grid of test inputs. Checked with the pydantic version recorded in the golden file.
- Limitation
- Whether a request is clear is decided by a model, not by this stage.
- Source
- src/schemas.py, model_b_approach/pipeline.py
- What it does
- A model call that decides whether the request has enough usable information. It returns a status, a canonical intent, the usable facts and, when needed, a clarifying question.
- Why it exists
- Ask first instead of guessing.
- In and out
- A valid request. Ready with a canonical intent and usable facts, or clarify with a question.
- Decision
- Make clarification a normal outcome of the pipeline.
- What it does not cover
- It is a model, so it can miss: in one recorded scenario a missing extension length did not trigger a question.
- Evidence
Partially verifiedTen scenarios, one run: the gated pipeline asked first in 2 of the 3 scenarios that expected a question. Judge and generator models were not recorded; the result was not reproduced.
- Limitation
- Not run in this page. Outcomes are recorded, and any edit retires them.
- Source
- model_b_approach/input_checker.py
Acceptance rule
- What it does
- Compares the verify score of a draft answer with the acceptance threshold. At or above the threshold the answer is accepted and repair is skipped.
- Why it exists
- Acceptance is an explicit rule with a number, not an assumption.
- In and out
- A verify score from 0 to 1 and the threshold (0.72 by default). Accepted, or one repair attempt.
- Decision
- Make acceptance a threshold that a request can override.
- What it does not cover
- The verify score is only as meaningful as the verifier and the generator; the committed runs use a mock generator.
- Evidence
VerifiedCommitted runs record the final score, whether repair ran and whether the answer was accepted. Mock generator; scores describe harness behavior, not answer quality.
- Limitation
- No claim is made about answer quality or real model behavior.
- Source
- rag_papers/retrieval/router_dag.py, run_plan and exec_repair
- What it does
- Makes one more pass with tighter settings: more weight on BM25 (0.5 to 0.7) and less on vectors (0.5 to 0.3), a stricter prune, a constrained template and a lower temperature (0.2 to 0.1), then verifies again.
- Why it exists
- A bounded second attempt has a fixed cost, and acceptance is decided by the same rule.
- In and out
- A draft below the threshold. A new draft and a new verify score.
- Decision
- One repair only.
- What it does not cover
- After the single attempt the answer can still be below the threshold and is then returned as not accepted.
- Evidence
Partially verifiedIn the committed runs that produced scores, every query that used repair ended below 0.72 and was not accepted. Mock generator; scores describe harness behavior.
- Limitation
- Whether a repair improves an answer is not claimed.
- Source
- rag_papers/retrieval/router_dag.py, exec_repair
- Start: a car with 3 seats1 of 3 taken
- A request for 2 seats3 of 3: accepted
- 2 more seats requestedwould be 5 of 3: refused car stays 3 of 3 · HTTP 409
- Back at 1 of 3: two requests for 2 seats at the same momentrequest 1 reads “1 of 3”: fits, so it booksrequest 2 reads “1 of 3”: fits, so it books5 of 3 observed
Engineering mechanism
Every status change goes through one transition table (REQUESTED, MATCHED, DRIVER_ARRIVED, STARTED, COMPLETED). The capacity guard runs inside the accept transaction and adds up the seats of every active pool of the vehicle; a refusal is HTTP 409. Two simultaneous accepts can both pass because the check reads, then writes (seen under an in-memory database mock).
Another request for 2 seats
Refused3 + 2 is more than 3, so the car stays at 3 of 3. HTTP 409
Where the check sits in a request's life
- REQUESTED
- MATCHEDaccept: capacity guard
- DRIVER_ARRIVED
- STARTED
- COMPLETED
Every change goes through one transition table. Cancellation is allowed before the ride starts.
Count the whole car, not each group
- each group on its ownlooks finepasses
- the whole car added up3 + 2 is more than 3refused
Before and after the seat check was added
Two requests at the same moment
Each request checks the seats, sees room, then books. Neither sees the other, so both get in.
accept 1reads 1 of 3passeswrites
accept 2reads 1 of 3passeswrites
“I need an email about my job assessment.”
Engineering mechanism
Request validation in code (intent not empty, at least one non-empty fact, tone from a fixed list), then the input checker, a model call that decides whether there is enough usable information (not run in this page), then a clarify or generate branch. An evaluation runner scores both strategies on ten recorded scenarios. This is scenario 3; the scores are recorded judge outputs, and why the checker decided as it did was not saved.
First check: does the request have the basics?
- it says what the email is for
- it gives at least one fact no facts given
- the tone is one of professional, polite, formal, friendly, empathetic, urgent but respectful, casual but respectful
Shown: a request with every fact removed. It stops here, before the AI is asked to write anything.
Technical message
List should have at least 1 item after validation, not 0
Is there enough to write from?
a request with the basics
decides if there is enough to write frominput checker: a model call, not run in this page
asks a question firstrecorded for scenario 3: writing stops until the request is clearer
writes the emailthe other branch
Tested on 10 recorded scenarios
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- 10
9 of 10: asked or wrote, as expected. Average score 4.7 out of 5, against 4.2 for a single prompt. One run, scored by a model.
One unclear request still got through
- request
- Request a deadline extension for my final project.
- should have
- asked first
- actually
- wrote the email
Engineering mechanism
The route is retrieve, rerank, prune, contextualize, generate, verify, and repair below the threshold. The verify score (0 to 1) is rule-based: length 20%, relevance 30% (how many of the question's words the answer uses), coherence 25%, completeness 15% and repetition 10%. It does not compare the answer with the retrieved passages. The default accept threshold is 0.72. The first draft's score was not recorded; q3's final score after repair was 0.67. Drafts came from a mock generator.
Is the draft good enough?
Below the required qualityThe draft scored 0.67; the requirement is 0.72.
A second attempt, once
- first draft
- below the requirement
- one repair attempt
- checked again
the system does not just return its first answer
Engineering route
- retrieve
- rerank
- prune
- contextualize
- generate
- verify
Below the threshold: one repair attempt, then verify again.
Repair attempt: bm25 0.5 to 0.7, vector 0.5 to 0.3, prune overlap 1 to 2, constrained template, temperature 0.2 to 0.1.
What happened to four recorded questions
| q1 | 0.82 | accepted | |
|---|---|---|---|
| q5 | 0.72 | accepted | |
| q3 | 0.67 | tried again, still below | |
| q4 | 0.655 | tried again, still below |
A repair is an attempt, not a guarantee
Engineering mechanism
The route is OCR, a lab-report check, metadata, the value parser and row omission. The parser accepts a plain number, a qualified number (such as <8.5) or a scientific value, and leaves out any row whose value matches no known form, keeping the raw line for rows that parse. A misread digit that still forms a valid number passes, and there is no confidence measure. The README records <8.5 read in place of <0.5.
How a value is read
- test
- CRP
- value
- <0.5
- unit
- mg/dL
- range
- <1.0
- flag
- N/A
The value is accepted only if it has a shape the parser knows: a plain number, a qualified number or a scientific value. Nothing else.
A reading it cannot make sense of
- test
- Cell Count
- value
- 1.2 x 1043
- unit
- 10^3/µL
- range
- 1.0 - 2.0
- flag
- N/A
row left outIt is not guessed and not kept.
One misread is caught, one is not
- 1.2 x 1043does not look like a numberrow left out
- <8.5looks like a valid numberrow kept
It looks valid, but it is wrong
- test
- RP
- value
- <8.5
- unit
- mg/dL
- range
- <1.0
- flag
- N/A
parsedlooks like a valid value
should read<0.5
was read as<8.5
keptNothing flags it, so a reader would never know.
Other systemsselect one to switch
Ride-pooling lifecycle with guarded transitionsoverbooking
How do you stop a car being overbooked?
A ride-sharing backend that refuses a booking when the car would end up with more riders than seats. It does not yet hold when two bookings arrive at the same instant.
My part: the ride service, the request steps and the seat check. Sole contributor in the commit history (all 125 commits).
What goes wrong: Ride-pooling lifecycle with guarded transitions
- Goes wrong
- Several riders can ask for seats in the same vehicle, and each request looks fine on its own.
- Why it matters
- If every request is accepted separately, the car can end up with more riders than it has seats.
- What it does
- It adds up the seats already taken across the whole vehicle and refuses a request that would not fit.
In engineering terms
A shared-ride service where a request moves only along a transition table, with a capacity check across a vehicle's pools that holds for sequential accepts; concurrent accepts remain a known gap.
Pooling riders into shared trips under constraints while several actors change state at once, with an auditable history.
TypeScript, Next.js, Express, PostgreSQL
How it responds: Ride-pooling lifecycle with guarded transitions
Select a step to see what it does and why it exists.
A rider asks for seats Request
A passenger creates a request with a seat count, and the server computes an integer-paisa fare.
Decision. Integer paisa for money.
Every request follows fixed steps Transition table
Every status change goes through one table: REQUESTED, MATCHED, DRIVER_ARRIVED, STARTED, COMPLETED, with cancellation allowed before the ride starts.
Decision. One transition table for every status change.
more about Transition table
- Why it exists
- So no route can skip a state.
- Evidence
Verified 71 written test cases exist in the repository. They were counted, not run for this page.
- Source
- apps/api/src/common/rideLifecycle.ts
A driver accepts the request Conditional accept
A driver accepts a request. A conditional update succeeds only if the request is still REQUESTED.
Decision. Conditional (compare-and-set) updates for single-row races.
more about Conditional accept
- Limitation
- It covers a single row, not the vehicle's total across pools.
Riders are grouped into one trip Pool matching
The service joins an existing pool when routes are compatible, or founds a new pool.
more about Pool matching
- Why it exists
- Pools form by route compatibility, within pickup and destination distance limits.
Blocks overbooking Capacity guard
Adds up the seats committed across every active pool of the vehicle (open or locked) and compares that total plus the new request with the vehicle's capacity.
Decision. Guard the vehicle's total, not each pool.
more about Capacity guard
- Why it exists
- A per-pool check passed while the vehicle as a whole was over capacity, so the guard was moved to the vehicle level.
- In and out
- Requested seats (1 to 6) and the vehicle's active pools. Accept, or HTTP 409 and a conflict status event.
- Failure
- Two accepts at the same moment can both read a total under capacity before either writes.
- Control
- The accept transaction refuses when the total plus the request exceeds capacity.
- Evidence
Verified Recorded before and after the guard commit: 5 of 3 seats without it, 3 of 3 with it. In-memory database mock, not PostgreSQL.
Partially verified Two accepts at the same moment both passed the guard under the mock, reaching 5 of 3. PostgreSQL behavior was not run.
- Limitation
- Concurrent accepts across pools are not covered by this guard.
- Source
- apps/api/src/modules/rides/rides.service.ts, accept transaction
Every change is logged Status events
Every transition leaves a status event, and a conflict outcome is written as an event after the rollback.
Decision. Success events committed with state changes, and conflict events written after a rollback.
Decision: Ride-pooling lifecycle with guarded transitions
Count seats for the whole car, not for each group of riders, at the moment a driver accepts a request.
In engineering terms
Guard the vehicle's total seats across pools, not each pool, inside the accept transaction.
What happened: Ride-pooling lifecycle with guarded transitions
Recorded test run: without the check a 3-seat car held 5 riders; with it, 3 of 3. This ran against a stand-in (mock) database, not a real one.
Evidence detail Checked by me
71 written test cases exist in the repository. They were counted, not run for this page.
Where it still fails: Ride-pooling lifecycle with guarded transitions
Two requests accepted at the same moment can both pass the check, which would leave the car at 5 of 3. Seen only against a stand-in (mock) database; a real database was not run.
In engineering terms
Concurrent accepts across pools are not covered by the guard. Observed under an in-memory mock; PostgreSQL was not run.
Case study →: Ride-pooling lifecycle with guarded transitionsWatch it run (Ride-pooling lifecycle with guarded transitions)
Two-stage email drafting pipelineunclear requests
When should an AI not start writing?
An AI email writer that first checks whether the request makes sense, and asks a question instead of writing when it does not.
My part: the design of both approaches, the request checker, the scoring framework, the ten test scenarios and the demo app. Sole contributor in the commit history.
What goes wrong: Two-stage email drafting pipeline
- Goes wrong
- A single prompt writes a confident email even when the request is vague, contradictory or missing the facts it needs.
- Why it matters
- A confident email looks finished, so nobody notices when it says the wrong thing.
- What it does
- It checks the request first, and asks a question instead of writing when there is not enough to go on.
In engineering terms
Compares a single prompt with a pipeline that checks inputs before it generates, and scores both against a fixed set of scenarios.
A single prompt writes a confident email even when the request is unclear, contradictory, padded with irrelevant facts or contains instruction-like text.
Python, Streamlit, LLM provider APIs, Pydantic
How it responds: Two-stage email drafting pipeline
Select a step to see what it does and why it exists.
Checks the request has the basics Request validation
Checks the request in code before any model is called: the intent is not empty, there is at least one non-empty fact, and the tone is one of the allowed values.
Decision. Validate in code first.
more about Request validation
- Why it exists
- Cheap, deterministic failures should never reach a model.
- In and out
- Intent, key facts and tone. A valid request, or an error that stops the run.
- Failure
- Code can only check shape. It cannot tell whether the request is clear enough.
- Control
- A validation error is raised before the input checker runs.
- Evidence
Verified The messages and rules shown here match the original Python on a grid of test inputs. Checked with the pydantic version recorded in the golden file.
- Limitation
- Whether a request is clear is decided by a model, not by this stage.
- Source
- src/schemas.py, model_b_approach/pipeline.py
Decides if there is enough to write from Input checker
A model call that decides whether the request has enough usable information. It returns a status, a canonical intent, the usable facts and, when needed, a clarifying question.
Decision. Make clarification a normal outcome of the pipeline.
more about Input checker
- Why it exists
- Ask first instead of guessing.
- In and out
- A valid request. Ready with a canonical intent and usable facts, or clarify with a question.
- Failure
- It is a model, so it can miss: in one recorded scenario a missing extension length did not trigger a question.
- Control
- A clarify status returns the question and no email is generated.
- Evidence
Partially verified Ten scenarios, one run: the gated pipeline asked first in 2 of the 3 scenarios that expected a question. Judge and generator models were not recorded; the result was not reproduced.
- Limitation
- Not run in this page. Outcomes are recorded, and any edit retires them.
- Source
- model_b_approach/input_checker.py
Asks a question, or writes Clarify or generate
On a clarify status the pipeline returns the question and no email. Otherwise the generator writes the email.
Decision. Make clarification a normal outcome of the pipeline.
more about Clarify or generate
- Source
- model_b_approach/pipeline.py
Scores the results on test scenarios Evaluation runner
A separate runner sends scenarios through both strategies and scores the results with three judge metrics and structural checks. It is not part of the pipeline.
more about Evaluation runner
- Evidence
Partially verified The committed summary averages are 4.7 for the gated pipeline and 4.2 for the baseline, on a scale of 1 to 5. Ten scenarios, one run; the judge and generator models are not recorded.
- Source
- evaluate.py
Decision: Two-stage email drafting pipeline
Check the request with ordinary code first, then let a model judge whether there is enough to write from.
In engineering terms
Validate in code first, then let a model decide whether to ask a question before any email is written.
What happened: Two-stage email drafting pipeline
Tested on 10 recorded scenarios: the pipeline averaged 4.7 out of 5 and a single prompt 4.2. One run, scored by a model that was not recorded.
Evidence detail Checked by me
Ten test scenarios are saved in the repository, with the recorded outputs and the scores a judge model gave them.
Where it still fails: Two-stage email drafting pipeline
One unclear request still got through. For "Request a deadline extension for my final project" it wrote an email when it should have asked.
In engineering terms
Ten scenarios and one run. The checker is a model: in one recorded scenario it did not ask when it should have.
Case study →: Two-stage email drafting pipelineWatch it run (Two-stage email drafting pipeline)
Retrieval over research papers with verify and repairweak answers
When should an AI trust its first answer?
An AI that answers questions about research papers, scores its own draft for quality, and gets one retry when the draft falls short.
My part: the whole system: reading the papers, finding sources, writing and checking answers, and the code that tests it. Sole contributor in the commit history (all six commits).
What goes wrong: Retrieval over research papers with verify and repair
- Goes wrong
- A first draft can be too short, drift away from the question, read badly or stop mid-thought, and nothing would tell the reader.
- Why it matters
- An answer returned without any check is only as good as the first attempt.
- What it does
- It scores each draft with simple quality rules. Below the required quality, it makes one more attempt with tighter settings and scores again.
In engineering terms
A plan-driven pipeline that retrieves, scores the draft answer with a rule-based quality check and makes one repair attempt when the score is below a threshold.
Answering questions over research papers, including tables, with answers that are checked rather than trusted.
Python, FastAPI, DuckDB, sentence-transformers
How it responds: Retrieval over research papers with verify and repair
Select a step to see what it does and why it exists.
Works out what the question needs Query plan
A plan is chosen from the question's intent and then executed.
Decision. A plan keyed on query intent instead of one fixed chain of steps.
Finds supporting passages Hybrid retrieval
Two indexes, BM25 and an in-memory embedding store using cosine similarity, behind an ensemble retriever.
Decision. Hybrid lexical and vector retrieval.
Drops sentences that do not help Sentence-level prune
Irrelevant sentences are pruned before the context is built.
Decision. Sentence-level pruning before contextualization, so the generator sees less irrelevant text.
Checks the answer is good enough Acceptance rule
Compares the verify score of a draft answer with the acceptance threshold. At or above the threshold the answer is accepted and repair is skipped.
Decision. Make acceptance a threshold that a request can override.
more about Acceptance rule
- Why it exists
- Acceptance is an explicit rule with a number, not an assumption.
- In and out
- A verify score from 0 to 1 and the threshold (0.72 by default). Accepted, or one repair attempt.
- Failure
- The verify score is only as meaningful as the verifier and the generator; the committed runs use a mock generator.
- Control
- A score at or above the threshold is accepted; below it, one repair runs.
- Evidence
Verified Committed runs record the final score, whether repair ran and whether the answer was accepted. Mock generator; scores describe harness behavior, not answer quality.
- Limitation
- No claim is made about answer quality or real model behavior.
- Source
- rag_papers/retrieval/router_dag.py, run_plan and exec_repair
One more attempt if it falls short Repair step
Makes one more pass with tighter settings: more weight on BM25 (0.5 to 0.7) and less on vectors (0.5 to 0.3), a stricter prune, a constrained template and a lower temperature (0.2 to 0.1), then verifies again.
Decision. One repair only.
more about Repair step
- Why it exists
- A bounded second attempt has a fixed cost, and acceptance is decided by the same rule.
- In and out
- A draft below the threshold. A new draft and a new verify score.
- Failure
- After the single attempt the answer can still be below the threshold and is then returned as not accepted.
- Control
- The maximum number of repairs is 1.
- Evidence
Partially verified In the committed runs that produced scores, every query that used repair ended below 0.72 and was not accepted. Mock generator; scores describe harness behavior.
- Limitation
- Whether a repair improves an answer is not claimed.
- Source
- rag_papers/retrieval/router_dag.py, exec_repair
Decision: Retrieval over research papers with verify and repair
Do not return the first draft. Score it, accept it at the required quality, and allow one bounded repair attempt below it.
In engineering terms
Score each draft with the rule-based verifier, accept at a threshold, and make one bounded repair attempt below it.
What happened: Retrieval over research papers with verify and repair
Four recorded queries: two accepted on the first draft, two tried again and still below the requirement. The drafts came from a stand-in (mock) generator.
Evidence detail Checked by me
Saved evaluation runs record which path each question took through the pipeline.
Where it still fails: Retrieval over research papers with verify and repair
Repair is an attempt, not a guarantee: both repaired answers stayed below the requirement. The drafts came from a stand-in (mock) generator and the score is a simple rule-based quality check, so it shows how the checking behaves, not how good or correct the answers are.
In engineering terms
The committed runs use a mock generator, so scores describe harness behavior and not answer quality. Repair is not a guarantee.
Case study →: Retrieval over research papers with verify and repairWatch it run (Retrieval over research papers with verify and repair)
Speech and report extraction that does not invent valuesmisread values
What if a reading looks valid but is not?
A system that transcribes speech and pulls structured data out of photographed lab reports, leaving out any value it cannot read instead of guessing.
My part: the whole service: its web endpoints, the code that reads and sorts report text, and the tests and sample data. Sole contributor in the commit history (all 51 commits).
What goes wrong: Speech and report extraction that does not invent values
- Goes wrong
- Reading text from a photographed lab report sometimes misreads a value, and a misread can still look like a perfectly good number.
- Why it matters
- A wrong lab value that looks valid is worse than a missing one, because nothing warns the reader.
- What it does
- It turns report text into structured fields and leaves out values it cannot read. The page also shows the case it cannot catch.
In engineering terms
Transcribes speech and extracts structured rows from photographed lab reports, omitting rows whose values cannot be parsed and returning zero rows for documents that are not lab reports.
Turning messy audio and photographed documents into structured data without making values up.
Python, FastAPI, faster-whisper, Tesseract
How it responds: Speech and report extraction that does not invent values
Select a step to see what it does and why it exists.
Reads the text from the photo OCR
A provider reads a photographed report into text lines. A mock provider is the default and Tesseract is the real one.
Decision. Provider adapters with mock defaults.
more about OCR
- Limitation
- OCR degrades on rotated or angled images, and raw outputs are not committed.
Checks it is a lab report Lab report check
A conservative detector decides whether the text looks like a lab report. If not, the type is unknown and zero rows are returned.
Decision. Conservative non-lab detection.
more about Lab report check
- Evidence
Partially verified A receipt returned the type unknown and 0 rows. Tested on one image; self-reported.
Picks out patient, date and lab Metadata
Reads patient, date and lab fields where present. With the header cropped out, metadata is null.
more about Metadata
- Evidence
Partially verified With the header cropped out, metadata is null and 4 rows are recovered. Self-reported README tables.
Turns report text into fields Value parser
Reads the value column as a plain number, a qualified number, a range or a scientific value. Anything that matches none of those forms returns no value.
Decision. Reject malformed values, and do not correct OCR using knowledge of the source document.
more about Value parser
- Why it exists
- A value that does not match a known form is not repaired or guessed.
- In and out
- The value text of one row. A value with its kind, operator, numbers and the raw text, or none.
- Failure
- A wrong value that is still well formed parses as valid, for example <8.5 read in place of <0.5.
- Control
- No match returns none, and the row is omitted.
- Evidence
Partially verified The README records Tesseract reading CRP <0.5 as RP <8.5 and 1.2 x 10^3 as 1.2 x 1043. Self-reported evaluation; raw OCR outputs are not committed.
- Limitation
- There is no confidence measure, so the parser cannot detect a misread that is still a valid number.
- Source
- app/services/value_normalizer.py
Leaves out values it cannot read Row omission
Drops a row whose value does not parse. A row that does parse keeps its raw line.
Decision. Omit rather than guess.
more about Row omission
- Why it exists
- A made-up value is worse than a missing row.
- In and out
- A parsed value, or none. A result row with its raw line, or no row.
- Failure
- An omitted row is silent: the response does not list what was dropped.
- Control
- A value of none returns no row.
- Evidence
Partially verified Recorded examples: with the header cropped out, 4 rows are recovered; a receipt returns the type unknown and 0 rows. Self-reported README tables.
- Limitation
- The OCR text of dropped rows is not returned.
- Source
- app/services/report_parser.py
Decision: Speech and report extraction that does not invent values
Accept only value shapes it knows. Leave out any row that does not match, and keep the raw line for rows that do.
In engineering terms
Reject a value that does not match a known form, omit that row, and keep the raw line on rows that parse.
What happened: Speech and report extraction that does not invent values
Two recorded misreads: one is caught and its row left out, the other looks valid and is kept. Test recordings and synthetic report images are saved in the repository.
Evidence detail Checked by me
Test recordings and synthetic report images are saved in the repository.
Where it still fails: Speech and report extraction that does not invent values
It cannot tell a wrong value that looks valid. A reading of <8.5 in place of <0.5 is parsed and kept, and nothing flags it. There is no confidence measure.
In engineering terms
A misread that is still a valid number parses as valid. There is no confidence measure.
Case study →: Speech and report extraction that does not invent valuesWatch it run (Speech and report extraction that does not invent values)
What you tested
Nothing tested yet. Operate any system above and this page writes down what you tried, what held and what did not.
copied