INSIGHT · AI GOVERNANCE & COMPLIANCE

AI reads the text on a drawing well.
It still miscounts the doors.

Published research says it depends on what is being read. In AECV-Bench (arXiv preprint, January 2026), the best of ten multimodal AI models reached an exact-match accuracy of 0.39 for counting doors and 0.34 for windows across 120 floor plans, while top models scored 0.7 to 0.95 on questions that read text from drawings. Models trained for floor plans are scored on their own datasets with other metrics, and lose accuracy when the drawing style changes. None of the datasets is described by its authors as Australian residential drawings.

Close-up of a printed architectural floor plan

Analysed 9 October 2026 · AECV-Bench (arXiv 2601.04819, v1 8 January 2026) · FloorPlanCAD (ICCV 2021, arXiv 2105.07147 v2) · CubiCasa5K (arXiv 1904.01920, 2019) · Parsing Floor Plan Images (MVA 2017) · RMIT framework paper (arXiv 2607.00015, 26 May 2026) · all papers read 9 October 2026.

Key takeaways

  • Text is the strong case: On AECV-Bench's drawing questions, text extraction is the strongest category for top models, at 0.7 to 0.95 accuracy, on the text-extraction questions within a set of 192 question-answer pairs on 21 drawings, each scored correct or incorrect.
  • Door and window counts are the weak case: On 120 floor plans, the best exact-match accuracy any of the ten models reached was 0.39 for doors and 0.34 for windows. Bedrooms (up to 0.91) and toilets (up to 0.82) scored higher, which the authors attribute to those rooms usually carrying a text label.
  • Purpose-trained models are measured on their own data: On FloorPlanCAD's vector CAD test drawings (per-class results reported on the paper's 30-class V1 dataset), the authors' model reached a panoptic quality of 0.763 for single doors and 0.459 for windows. A wall model trained only on European plans reached 76.1 mean IoU on Japanese real-estate plans, against 90.5 for a model trained on both.
  • Mark-ups and revisions are not measured: None of the papers we read reports accuracy on hand-annotated drawings or on tracking a change between revisions. AECV-Bench states that each item is a single image with no cross-referencing to other sheets.
  • We would store where each count came from: A system-design recommendation: for any quantity drawn from a drawing, record the sheet and revision, whether the file was vector or scanned, whether an agent or a person produced the count, and who confirmed it.

01

Five papers, five meanings
of "accuracy".

Each study measures a different task on a different set of drawings with a different metric. A number from one cannot be compared with a number from another, and none of them is a measure of a product.

AECV-Bench and the RMIT paper are arXiv preprints; their arXiv pages showed no journal reference as read on 9 October 2026. FloorPlanCAD appeared at ICCV 2021 and Parsing Floor Plan Images at MVA 2017.

PaperTask measuredDrawingsMetric
AECV-Bench (arXiv preprint, 2026)General multimodal models counting doors, windows, bedrooms, toilets; answering questions about drawings120 floor plans from CubiCasa5K, CVC-FP and public sources; 192 questions on 21 drawingsExact-match accuracy per class; mean absolute percentage error (MAPE); binary correctness judged by a language model with human adjudication
FloorPlanCAD (ICCV 2021)Spotting 30 to 35 symbol and element classes in CAD drawingsOver 15,000 vector CAD floor plans, residential to commercial; test set 5,502Panoptic quality (PQ), weighted F1, mean average precision (mAP)
CubiCasa5K (2019)Segmenting rooms and icons such as doors and windows5,000 raster floor plans, mostly Finnish; test set 400Pixel accuracy and mean intersection over union (IoU)
Parsing Floor Plan Images (MVA 2017)Segmenting walls; detecting objects; reading room sizes with OCRR-FP: 500 Japanese real-estate floor plans; CVC-FP: 122 European plansMean IoU; average precision; share of room-size labels read
RMIT framework (arXiv preprint, 2026)A proposed compliance-checking framework for multi-apartment plansAustralian apartment design policies as contextNo accuracy figures reported

02

Symbols:
doors and windows.

Doors and windows are drawn as line-work conventions, not words. Each paper measures them with a different test, and the results range widely.

AECV-Bench counting rules are specific: double-leaf doors count as two, sliding openings and French doors count as windows, adjacent windows without a gap count as one. A count under different rules is a different number.

Study and metricDoorsWindows
AECV-Bench, best exact-match accuracy of ten models, 120 plans0.390.34
AECV-Bench, lowest mean absolute percentage error of ten models15.0%20.4%
FloorPlanCAD, panoptic quality, authors' model, V1 test set0.763 (single door), 0.748 (double door)0.459 (window), 0.154 (bay window)
CubiCasa5K, test IoU, raw segmentation53.666.8
CubiCasa5K, test IoU, after converting to polygons41.240.9
Parsing Floor Plan Images, average precision on 25 test images96.0% (doors), 35.9% (sliding doors)Not reported
  • Exact match is a strict test: AECV-Bench scores a plan as correct only when the predicted count equals the hand-labelled count. The MAPE figures show how far off the counts were in proportion when they missed.
  • Converting to usable shapes loses accuracy: CubiCasa5K reports lower scores after segmentations are turned into polygons, and explains why: "if wall or icon junctions are missed or are not correctly located the polygons can not be created regardless the quality of the segmentation."

03

Text:
labels, scales and room sizes.

Reading text printed on a drawing, such as room names, area values, sheet numbers, scales and dates, is where general multimodal models score best. AECV-Bench reports text extraction at 0.7 to 0.95 for top models, spatial reasoning at 0.6 to high 0.7, and instance counting as "the most error-prone component of the QA suite".

  • Labelled rooms are counted better than drawn symbols: Best exact-match accuracy on AECV-Bench was 0.91 for bedrooms and 0.82 for toilets. The authors attribute this to rooms that are "frequently explicitly labeled", letting models "fall back on OCR and tag-based reasoning".
  • Room sizes from an older pipeline: Parsing Floor Plan Images labelled 20 R-FP images with text locations and values and reports that optical character recognition (OCR) detected and recognised 74.1% of the room size annotations.
  • The same models, the same drawings: AECV-Bench notes that counting errors "persist even when the same models perform well on OCR for the very same drawings", locating the difficulty in "understanding the graphical language of AEC plans".

04

What fails, as the papers
themselves describe it.

The authors' own descriptions of failure cluster around drawing style, density and the format of the file.

We found no paper among those read that reports accuracy on drawings carrying hand mark-ups, or on detecting what changed between two revisions of a sheet. That is a gap in the evidence, not a finding that it fails.

  • Misread conventions: AECV-Bench: "Models frequently misinterpret door swings, confuse windows with openings or façade elements, hallucinate fixtures, or miss instances in dense regions."
  • A new drawing style: Parsing Floor Plan Images trained a wall model on one dataset and tested it on the other. Trained on CVC-FP and tested on R-FP it reached 76.1 mean IoU; trained on both, 90.5. The authors attribute the gap "largely" to "differences in the drawing styles".
  • Clean test sets, messy real sets: The RMIT paper (2026) notes that "many existing methods are evaluated on relatively clean datasets, whereas real-world architectural drawings frequently contain dense annotations, inconsistent symbols, overlapping text, incomplete labels, and multi-scale details."
  • Poor scans: The same paper states that plans with "severe noise, low resolution, incomplete annotations, or ambiguous graphical conventions may not support reliable OCR, symbol recognition, or structural parsing, and may therefore require manual review".
  • One sheet at a time: AECV-Bench's limits state that each item "presents a single drawing image without cross-referencing to other sheets, details, or specifications". Tracing a callout to a detail sheet, or a change from one revision to the next, is not tested.

05

Vector or scan:
the file changes the problem.

A residential drawing set can reach a builder as a CAD export, a vector PDF, or a scan of a printed sheet. The research is split the same way.

Which kind of file a quantity came from is part of how far it can be trusted. A figure measured on vector CAD drawings says little about a scanned PDF of the same house.

  • Vector: FloorPlanCAD's drawings "are all represented as vector graphics", with line-grained annotations. Its scores describe reading geometry that already exists as lines.
  • Raster: In CubiCasa5K, "a sample input is always a raster scan (usually a scanned copy)". AECV-Bench evaluates "rasterized PNG images of drawings rather than native CAD or BIM formats".
  • Scale has to be recovered: The RMIT framework proposes reading the PDF's physical paper size and rasterising at a known resolution to map pixels to real-world units. It is a proposed method; the paper reports no measured accuracy for it.

06

Why a person confirms,
in the researchers' words.

The AECV-Bench authors draw the operational line themselves. Symbol-grounded counting "remains unreliable for automation without tight human supervision or custom trained models. This is especially important for workflows where counts flow directly into estimates, schedules, or compliance checks." They recommend counting features be "deployed as assistive tools that surface candidates and uncertainty, rather than as autonomous extractors".

  • The RMIT authors reach the same design: Their framework uses confidence scores to decide "when to escalate a case to a human reviewer", serving "as a first-pass screening tool" rather than "replacing human oversight".
  • Our own research listing: Zhang, Y., Chang, R., Fang, X., & Jiang, W. (2026). AI-Assisted Decision Support for Drawing-Based Residential Compliance Review: Integrating Large Language Models and Computer Vision. WSBE26, Melbourne. Our publications listing describes it as defining how drawings are read, "with the final compliance decision kept with a human reviewer". It is listed without accuracy figures.

What we would put in a system

Any quantity or dimension taken from a drawing, by a person or an agent, is only as good as the record of where it came from. For each extracted value, we would test for five fields.

Where agents help is narrow. A document intelligence agent can read the title block, schedules, room labels and dimensions from a drawing, propose counts of doors and windows with uncertainty marked, and draft the record; a person confirms each value before it is used in a quote, an order or a compliance submission. No agent decides whether a design complies, and no count is relied on unconfirmed.

FieldWhy it is load-bearing
Sheet number and revisionAECV-Bench tests one image at a time; a value without its revision cannot be checked against the next issue of the sheet.
File type: vector, vector PDF or scanThe research separates vector and raster inputs, and reports lower scores where drawing style or quality differs from the training data.
Element class and counting ruleWhether a double door is one or two, and whether a sliding door is a window, changes the number.
Produced by agent or person, with any flagDoor and window counts are the weakest class in AECV-Bench; a flagged agent count needs a different check from a text reading.
Confirmed by, and whenThe person who checked the value against the drawing before it entered a quote, order or programme.

Questions worth asking
of your own drawings

For an estimator: which quantities in your last quote came from counting symbols on a plan, and who checked them?

Door and window counts are the weakest class in AECV-Bench. A quote built on unconfirmed symbol counts carries the error rate of whichever method produced them.

For a builder: do your drawing sets arrive as CAD or vector PDF, or as scans?

Published accuracy figures come from either vector drawings or raster images, rarely both. Knowing which you hold tells you which body of research describes your case.

For a design or drafting consultant: do your plans label rooms in text, and are window and door types tagged?

AECV-Bench found labelled rooms counted far more reliably than drawn symbols. Text tags on openings give a reader, human or model, something firmer than line-work to rely on.

For a manufacturer quoting windows from plans: is each quoted opening tied to a sheet and revision?

No paper we read measures change detection between revisions. If the revision is not stored with the count, a superseded sheet is indistinguishable from the current one.

Questions people ask
about AI and construction drawings

Can AI read construction drawings accurately?

It depends on what is read. On AECV-Bench (arXiv preprint, 2026), top general multimodal models scored 0.7 to 0.95 accuracy on text-reading questions about drawings, but the best exact-match accuracy for counting doors on 120 floor plans was 0.39 and for windows 0.34.

How accurate is AI quantity takeoff from floor plans?

Published research reports element-level scores, not takeoff accuracy. FloorPlanCAD's model reached a panoptic quality of 0.561 across all classes on its vector CAD test drawings (V1 dataset, 30 classes), with 0.763 for single doors and 0.459 for windows. A separate 2017 study found a wall model scored lower on a drawing style it was not trained on.

Can AI extract dimensions and room sizes from plans?

Reading printed values is the strongest task in the research we read. An older study (MVA 2017) reports OCR recognising 74.1% of room size annotations on 20 labelled Japanese real-estate plans; AECV-Bench reports 0.7 to 0.95 for text extraction by top current models.

Does AI work on scanned PDFs of drawings?

AECV-Bench tests raster images and reports lower scores for counting symbols than for reading text; CubiCasa5K, also raster, reports test IoU of 53.6 for doors and 66.8 for windows. The RMIT paper (2026) states that plans with severe noise or low resolution may require manual review rather than automated reading.

What this analysis
does and does not show.

Evidence note

What it shows
What published research reports for AI reading floor plans and construction drawings: text extraction strongest, symbol counting weakest, purpose-trained models scoring higher on their own datasets and lower on new drawing styles, with each figure tied to its paper, dataset and metric as read on 9 October 2026.
Key facts quoted
AECV-Bench: best exact match 0.39 doors and 0.34 windows on 120 plans, 0.91 bedrooms and 0.82 toilets, text extraction 0.7 to 0.95 for top models; FloorPlanCAD: PQ 0.561 overall, 0.763 single door, 0.459 window (V1 dataset); CubiCasa5K: test IoU 53.6 door and 66.8 window, 41.2 and 40.9 after polygonisation; MVA 2017: 74.1% of room size labels read, wall mean IoU 76.1 cross-dataset against 90.5.
Derived
None. Selecting the best and lowest values across the ten AECV-Bench models is a reading of the authors' tables, not a calculation.
  • AECV-Bench and the RMIT paper are arXiv preprints with no journal reference as read. Three of AECV-Bench's four authors list a non-university affiliation; the fourth lists the University of New South Wales.
  • None of the datasets is described by its authors as Australian residential drawings. CubiCasa5K is mostly Finnish, R-FP Japanese and CVC-FP European; FloorPlanCAD spans residential to commercial buildings.
  • Scores are for the models and dates tested. They are not measurements of any commercial product, and we do not name the models tested.
  • No paper read reports accuracy on hand-annotated drawings, on multi-sheet sets, or on detecting changes between revisions.
  • Figures from different papers use different metrics and are not comparable with each other.
  • Nothing here is advice on whether any drawing, design or quantity complies with the National Construction Code or any standard.

Bring us one drawing set
and the quote built from it.

Tell us which quantities in the quote came from the drawings, whether the set was CAD, vector PDF or scanned, and who checked the counts. We will show you which of those values a model reads reliably, which it should only propose, and where the confirmation belongs.

General information about published research on automated drawing reading. It is not engineering, compliance or legal advice and does not assess any product, drawing or design. Figures are as reported in the papers listed, read on 9 October 2026, for the datasets and models they describe.

Send us the drawing set · Document intelligence

Sources

Suggested citation: AECV-Bench (arXiv 2601.04819), FloorPlanCAD (ICCV 2021), CubiCasa5K (arXiv 1904.01920), Dodge, Xu and Stenger, Parsing Floor Plan Images (MVA 2017), and Gautam et al. (arXiv 2607.00015), as read 9 October 2026. No figures are derived; all are as reported by the authors.