Edgent LLC · SDVOSB · UEI UZFLEDGS3PC5 · CAGE 1ZSR0
Sources Sought FA4890-260713 · AI for Aircraft Maintenance Assistance · HQ ACC AMIC
The maintenance record already contains the answer. Nothing can read it.
The notice asks for a commercial AI capability that cleanses and standardizes historical USAF aircraft maintenance data and provides prescriptive troubleshooting to maintainers, to inform an aircraft repair decision. A4A is that capability, runnable on this page: it reads the record (scanned paper and handwriting included), cleanses the free-text narrative maintainers actually write, assigns standardized Air Transport Association codes, and returns repair recommendations with a stated confidence and a citation to every record behind them.
All demonstration data on this page is synthetic and was authored by Edgent, with its accuracy measured and published, including where it falls short. What this demonstration is, and is not.
Working demonstration
From scanned record to repair decision, runnable here
- 1 Cleanse and standardize scan, read, code the record, handwriting included
- 2 Prescriptive troubleshooting history retrieved, repairs ranked, records cited
- 3 Feasibility and performance measured accuracy, including where it falls short
- 4 Operational integration governed runs, batch scale, search and locate
This is the capability the notice asks to evaluate, runnable: read the record, cleanse and standardize the narrative, assign an ATA chapter, retrieve the history that resembles it, and recommend repairs with cited records, with every measured number published below including the ones that fall short. Upload a scanned record and let the server read it, page by page, as a governed run; or choose one of the synthetic write-ups; or type your own in maintainer shorthand.
A maintainer does not type into a box; the record is a page. Upload one the way it actually arrives: a single scanned page or photo (PNG or JPEG), a PDF of any length, or a ZIP of scanned pages. Everything is read the same way. It is streamed to the server and opened there, never in your browser: a PDF or an archive is unpacked and rasterized with pdftoppm, and every page is queued as its own reading job, read and scored on its own. Nothing is unpacked or rendered in your browser, which is the whole point, so an encoding a browser renders blank is still read, and the thumbnails you get back, including the ones that failed, are the images the reader was actually handed. A page nobody looked at is how a blank render once passed for a reading. How it travels, so nothing fails mysteriously: an upload up to 64 MB is streamed through the server, and a larger upload goes in parts straight to storage from your browser, so a multi-gigabyte archive is not capped. Inside a ZIP, page images are read from .png, .jpg, .tif and .jp2 entries; the server skips anything else and the page says so rather than letting a missing entry pass for a read one. Everything is processed on Edgent's own hardware and stored as a governed record, so do not upload anything you would not publish.
Read → Extract
Whatever you uploaded, one page or a whole archive, was opened on the server, not in your browser. Each page above is queued, read and scored on its own, so one page failing leaves the rest intact, and the counts state how many succeeded and how many did not. The thumbnail under each entry is the image the reader was handed, fetched back out of storage; it is not a second rendering made here for display.
Try a real handwritten document
The draft SOO's hardest input is not typed shorthand, it is paper: scanned pages with handwritten entries, which is where Extract carries the risk. These are genuine handwritten government records in the public domain, from the Internet Archive's government-documents collection, and none of them has been through this pipeline. Open one, download a page image or the PDF, and drop the file into the upload control above. It runs the same Scan and Extract stages the capability statement measures: every page gets a checksum and a stored preview, a duplicate document is recognized by its content hash and not re-run, and a page the reader cannot read is routed to human review with zero confidence, never invented.
-
Walker v. Clements, judgment, 1747
A colonial court judgment, entirely handwritten in iron-gall ink on paper.
-
Shirkey v. Miller, summons, 1747
A handwritten summons for a colonial-era dispute, faded and creased the way real records arrive.
-
John Kinney, declaration of intent, 1852
A handwritten declaration of intent to become a United States citizen, in clerk's cursive.
-
Grove to Smith, land deed, 1851
A handwritten deed with metes and bounds, the dense legal cursive a records system actually has to read.
-
STAR GATE operational report 8908, 1989
A declassified tasking sheet whose substance is handwritten feedback on a printed form: mixed print and handwriting on one page.
A one-click "run this example" that fetches the file server-side is a scoped follow-up: platform rule is that orchestration runs in VForce Flow, so the fetch, hash check, ingest and indexing will ship as a Flow rather than a bespoke endpoint chain.
Shorthand, typos and non-standard acronyms are the point, write it the way it would really be typed. Maximum 400 characters.
Choose an example or type a write-up, then run the pipeline.
-
1 Ingest and cleanse
As written
Standardized
Rule-based and deterministic. Every substitution is attributable to a named rule, which is what keeps a maintenance record auditable.
-
2 Assign an ATA code
Chapter level only. Measured top-1 accuracy is 66.0% from the discrepancy alone, below the draft SOO's 90% threshold, see measured results.
-
3 Retrieve history and recommend
Confidence is the similarity-weighted share of the cited records whose repair held. Records where the fault returned count against it. Every recommendation opens to show the records behind it.
What this demonstration is, precisely
The eight examples replay real dense-retrieval output computed with BAAI/bge-m3 against the synthetic corpus. Free text you type is served by a lexical retriever running in this page, so the demonstration needs no backend and cannot become an open compute endpoint. Both engines' measured accuracy is published in measured results.
That split is a property of this web page, not of the product. A4A is server-side. It runs where the maintenance data already is, inside the customer's accredited boundary, with inference local to the deployment.
Much of that record is paper, scanned forms and logbook pages with handwritten entries. So this is not ETL, it is SET: Scan, Extract, Translate. It runs server-side, inside the customer's accredited boundary, where the maintenance data already lives.
A working demonstration, not a claim
This page carries a demonstration you can run yourself, on synthetic data, with its accuracy measured and published, including where it falls short. ATA chapter assignment measures 66.0% top-1 from the discrepancy alone. The draft SOO asks for 90%. That gap is stated here rather than hidden, along with what would have to change to close it.
Run the demonstration · Read the measured results · How it works
Read this first
All data on this page is synthetic and was authored by Edgent for this demonstration. It contains no IMDS or REMIS content, no real or invented tail numbers, no National Stock Numbers and no part numbers. ATA chapter numbers are the public ATA 100 chapter-level breakdown, used here only as classification labels.
This page describes a capability Edgent is offering to build. Edgent claims no past performance on this capability and none is implied. Every capability is marked shipped or design intent.
Feasibility and performance
What it scores, including where it falls short
Measured on 300 synthetic records built from 150 distinct base narratives across 15 ATA chapters, under grouped 5-fold cross-validation. Every number below is produced by the evaluation harness that ships with the demonstration. None is hand-entered.
The headline number, stated plainly
ATA chapter assignment measures 66.0% top-1 (198 of 300) from the discrepancy narrative alone, which is what is available when a write-up is first entered. It reaches 70.3% (205 of 300) when the corrective action is also present, which is only true after the repair.
The draft SOO asks for ≥ 90%. At write-up time this is 24.0 percentage points short. Edgent is not claiming the threshold is met.
ATA chapter assignment
| Method | Input | Top-1 |
|---|---|---|
| bge-m3, k-nearest neighbor | Raw | 42.0% |
| bge-m3, k-nearest neighbor | Cleansed | 46.7% |
| Lexical TF-IDF, k-nearest neighbor | Cleansed | 46.0% |
| bge-m3, nearest centroid | Raw | 51.0% |
| bge-m3, nearest centroid + ATA title anchor | Cleansed | 64.0% |
| bge-m3, nearest centroid | Cleansed, discrepancy only | 66.0% |
| bge-m3, nearest centroid | Cleansed, discrepancy + corrective action | 70.3% |
Two things this table shows that are worth stating rather than leaving to be noticed. Dense k-nearest-neighbor (46.7%) and the lexical TF-IDF baseline (46.0%) sit within one percentage point of each other on the same cleansed text, so at k-nearest-neighbor the embedding model buys almost nothing and the real gain comes from switching to a centroid classifier. That near-tie is itself a change: before the cleanser defect was repaired, the lexical baseline sat above dense k-nearest-neighbor. And anchoring on ATA chapter titles lowers accuracy (63.7% against 66.0%), so it is not recommended.
The finding that matters most
Holding the classifier and the split constant, cleansing raises accuracy from 51.0% to 66.0%, 15.0 percentage points. That is direct, measured evidence for the draft SOO's first objective: standardizing the narrative is not housekeeping, it is the thing that makes everything after it work.
Cleansing
| Measure | Result | What it means |
|---|---|---|
| Abbreviation expansion recall | 87.8% (1,101 / 1,254) | Shorthand correctly expanded. The cleanser's lexicon deliberately omits 23 of the 91 abbreviations present in the corpus, so this is a lower bound rather than a closed-book score. |
| Spelling repair rate | 70.4% (176 / 250) | Injected typos corrected. Short words are left alone deliberately: repairing them usually produces a different real word. |
| Token F1 against clean text | 0.955 | Overall agreement between the cleansed narrative and the authored clean form. |
| Vocabulary reduction | 23.2% (917 → 704 distinct forms) | The measure an analyst cares about. Fewer distinct surface forms means the same failure written five ways finally groups into one. |
Retrieval and confidence
The top retrieved record shares the correct ATA chapter 41.3% of the time. Confidence calibration is imperfect and is published as measured:
| Stated confidence | Records | Observed effective |
|---|---|---|
| 0–20% | 9 | 77.8% |
| 20–40% | 5 | 100.0% |
| 40–60% | 62 | 90.3% |
| 60–80% | 44 | 81.8% |
| 80–100% | 180 | 77.8% |
The bands do not track observed effectiveness monotonically, so the confidence figure is not yet usable as a probability. On a corpus this small that is expected, and calibration is a tuning task rather than an architectural one , but it is not done, and it is not claimed as done.
Why the number is what it is
This corpus is deliberately the hardest regime for the method. Each failure mode appears exactly once, as two degraded variants, and folds are grouped by base narrative, so at test time the system has never seen that failure mode in any form. A real IMDS or REMIS corpus contains the same failure mode hundreds or thousands of times, which is precisely the regime retrieval is strongest in.
Edgent is not claiming that closes the gap to 90%. What would have to be true: a labeled corpus of real records, sub-chapter ATA labels rather than chapter level, and an SME feedback loop on low-confidence assignments. That is what a pilot establishes, and it is why the measurement harness matters more than the current score.
Confusions concentrate where a maintainer would expect: engine fuel control against engine indicating, airframe fuel against engine fuel, hydraulic against landing gear. Best chapter is ATA 49 (APU) at 100%; worst is ATA 73 (Engine Fuel and Control) at 45%.
Real documents, measured
Government paper Edgent did not write, read by the same pipeline
Every page below is a real, publicly released government document, bundled with this page as a static image. Nothing is fetched at run time. Each card shows the page, what the reader returned from it, and, with the same weight, what that result is not evidence for.
These are a separate claim from the 66.0% ATA chapter accuracy measured on the synthetic corpus. None of these pages carries an ATA code, so none of them tests the classifier. They test reading.
Across the hand-transcribed subset, n = 6 pages, character error rate runs 0.19 to 0.28 on a typed deck log, 0.36 to 0.39 on a hand-filled service card, and 1.00 on free-form 18th century cursive. Of 225 pages attempted, 188 were read; the remaining 37 returned backend_unavailable and are excluded from the accuracy denominators rather than folded in as failures. The routing score sits at 0.91 to 0.94 across that entire error range, which is the same finding as on synthetic data: it is a routing disposition and it discriminates nothing about accuracy.
About this page
Published in support of Air Combat Command Sources Sought FA4890-260713, A4A – Aircraft Maintenance Assistance. A Sources Sought is market research; no award results from it.
All demonstration data is synthetic and authored by Edgent. Every measured number on this page is produced by the evaluation harness that ships with the demonstration, and the shortfall against the draft SOO's 90% threshold is stated rather than omitted.
This page makes no external requests. It loads no third-party scripts, fonts or analytics, and sets no tracking cookies. Fonts are local subsets.
Edgent LLC · 311 Rimrock Ct, Bastrop, TX 78602 · 480-414-5556 · edgentllc.com