Document Understanding Pipeline — Scanned Document Capture, Validation and Export
This process converts one scanned or PDF business document per run (invoices, receipts, W-9s, certificates of filing) into structured, checked data written out as spreadsheets. It reads the document with OCR, decides what type(s) of document it contains, pulls out the header fields and line-item tables, and applies arithmetic and confidence rules to decide whether the result can be trusted. Anything that fails a rule is routed to a human reviewer, and every correction the reviewer makes is fed back into the classification and extraction models so the next run needs less intervention. Successful documents are exported to Excel and the source work item is closed off.
At a glance
| Trigger | Two entry points exist and are near-identical: Main-ActionCenter.xaml (unattended, review tasks suspend the job) and Main-Attended.xaml (operator sits at a review station). Both read Data\Config.xlsx, then branch on the input flag in_UseQueue. If true, the next item is pulled from the work queue (a shared list of pending jobs held in UiPath Orchestrator) named in config key DocumentUnderstandingQueueName, in folder DocumentUnderstandingQueuePath; the document path sits on that item under the key named by TargetFileKey (expected to be TargetFile). If false, the path is passed in directly as in_TargetFile. In the attended version, if neither is supplied the operator gets a file picker (filter: all files). |
| Frequency | Not recorded in the source. Each run processes exactly one input file; no scheduler, dispatcher or looping wrapper appears in the outline. |
| Systems used | Excel (config and export workbooks); UiPath Orchestrator (work queue, stored settings/"assets", storage bucket, human-task catalog); OCR / document-AI services (digitisation, Intelligent Keyword Classifier, Form Extractor, regex extractor, DU project extractor, invoice/receipt ML trainers); the source PDF/image files; Orchestrator REST API (one-off setup only); human review interface (Action Center tasks, or attended Classification/Validation Station) |
| Inputs | Data\Config.xlsx (sheets Settings, Constants, AutocorrectOcrMistakes, InvoicePostProcessing, ReceiptPostProcessing, plus Assets); one source document per run; queue items; Orchestrator-stored settings listed on the Assets sheet; the project taxonomy (the catalogue of document types and their field names); the classifier learning file at ClassifierLearningFilePath; the OCR-correction mapping JSON in storage bucket StorageBucketName at path AutocorrectStorageBucketDirectoryPath |
| Outputs | One Excel workbook per extracted document in ExportsFolder, named <documentId>_<startPage>-<endPage>.xlsx, one worksheet per exported table; updated classifier learning file; extractor training data in InvoicesTrainingFolder / ReceiptsTrainingFolder; updated OCR-correction JSON in the storage bucket; queue item marked Successful, or Failed with a Business vs Application error type; progress markers on the queue item; a full log trail keyed by the run's logKey |
| Typical run | One input file → OCR → classification → per-document extraction → rule checks → optional human review → one workbook per document found in the file. Duration is dominated by any human review step, which can suspend an unattended job indefinitely. Not otherwise recorded in the source. |
| Owner | Not recorded in the source. |
Before you start
- Config file present and readable.
Data\Config.xlsxmust exist relative to the process folder and be openable. The developer's note is explicit: "The file should be accessible - no retry mechanisms are used here!" If it is missing, the run dies immediately with no friendly message, because every log and error text the process uses is itself defined in that file. - Orchestrator connectivity and folder access. The robot needs read access to the queue named in
DocumentUnderstandingQueueName(folderDocumentUnderstandingQueuePath), read access to every stored setting listed on theAssetssheet (each row names the setting and its Orchestrator folder path), and read/write access to the storage bucketStorageBucketName. - Human-task catalogue exists. Review tasks are created in the catalogue named by config key
ActionCatalog. Reviewers must have access to it, or unattended jobs will suspend and never resume. - OCR / document-AI credentials. The classifier and form extractor use config key
ApiKeywith endpointsClassificationEndpointandFormExtractorEndpoint. The DU project extractor references credential asset pathsAcademy/UiDemo_CredentialandShared/ACMECredential. The ML trainers useInvoicesEndpoint(defaulthttps://du.uipath.com/ie/invoices) andReceiptsEndpoint(defaulthttps://du.uipath.com/ie/receipts). - File-system paths writable.
ExportsFolder,InvoicesTrainingFolder,ReceiptsTrainingFolder, and the folder containingClassifierLearningFilePath(the process creates and deletes a.lock…copy alongside it). - First-time setup only.
QuickStart/InitializeOrchestrator/InitializeOrchestrator.xamlreads a setup spreadsheet and POSTs JSON payloads to the Orchestrator REST API to create the storage bucket, stored setting, task catalogue and queue that the main process depends on. Run this once per environment before the first production run; it is not part of the normal daily flow. - Know which entry point is deployed. Confirm with the process owner whether the unattended (
Main-ActionCenter) or attended (Main-Attended) variant is live, because the human-review experience differs completely.
Procedure
Stage 1 — Load configuration
- Open
Data\Config.xlsxand read, in order, the sheetsSettings,Constants,AutocorrectOcrMistakes,InvoicePostProcessingandReceiptPostProcessing. - From each sheet, take every row's
NameandValueinto a single in-memory settings list. If a row'sNameis blank, then skip that row (this lets the spreadsheet carry blank spacer rows). - Derive the default retry behaviour: maximum execution attempts and retry interval. Developer note: "out_MaxAttempts value cannot be lower than 1" — a config value of 0 is raised to 1.
- Generate a
logKeyfor the run. Every log line and every queue-item progress note from here on carries it, so a support person can pick one job out of a busy log.
Stage 2 — Initialise the run
- Load the project taxonomy — the catalogue of document types and the named fields belonging to each.
- Read the
Assetssheet. For each row, fetch the value of the named Orchestrator-stored setting (columns:Name,Asset,OrchestratorAssetFolderPath) with caching disabled, and write it into the settings list under the row'sName, overwriting any same-named value that came from the spreadsheet. This is how secrets and environment-specific paths stay out of the config file. - If a stored setting cannot be fetched, then: if a value for that
Namealready exists in the settings list, continue with it and move to the next row; if not, log an Error (ErrorMessage_AssetFailedToLoadplus the asset name) and abort the run. - The whole of stage 2 is wrapped in a retry loop. Developer note: "Although unusual, the connection to the Orchestrator might time-out. The retry mechanism is used to compensate for minor fluctuations in network stability."
- A placeholder exists here for process-specific initialisation ("Write your custom Initialization code here"); nothing is implemented in the outline.
Stage 3 — Get the work item (queue mode only)
- If
in_UseQueueis false, then skip this stage; the document path is already known fromin_TargetFileor the operator's file picker. - Clear the target-file value, then request the next item from the queue named in
DocumentUnderstandingQueueName, folderDocumentUnderstandingQueuePath. Retried on transient failure. - If no item is returned, then log a warning (
LogMessage_TransactionItemNotFound) and end the run. Nothing is processed and no failure status is recorded — an empty queue is a normal outcome, not an error. - If an item is returned but has no value under the
TargetFileKeykey, then log a warning (LogMessage_TargetFileMissingInTransactionItem) and end the run without processing. - Otherwise, take the document path from that key, and copy every field on the queue item into the settings list. This means the dispatcher that created the item can override any config setting on a per-document basis.
Stage 4 — Digitize (OCR)
- Log the start with the file name.
- Run OCR on the target file to produce the machine-readable text and a page/word layout model that later stages read positions from. Settings: OCR applied to PDFs automatically, checkbox detection on, parallelism unrestricted.
- Retries are governed by config key
MaxExecutionAttemptsDigitize, which overrides the run default. Developer warning: "An OCR engine that uses a paid license might incur extra costs when re-executing. This should be taken into consideration for the number of retries." - A placeholder exists for image pre-processing to improve OCR (the note gives greyscaling as an example); nothing is implemented.
Stage 5 — Classify and split
- Run the Intelligent Keyword Classifier against the digitised document, using endpoint
ClassificationEndpoint, API keyApiKey, and the learning file atClassifierLearningFilePath. Document splitting is on, so one input file can be resolved into several documents, each with its own page range. - Log every candidate document type with its confidence score, pipe-separated, so a reviewer can see how close the call was.
- Retries governed by
MaxExecutionAttemptsClassify. Same paid-licence caution as stage 4.
Stage 6 — Classification business-rule check
- If config
AlwaysValidateClassificationis true, then force the automatic-classification result to "not successful", so every document goes to a human regardless of confidence. Developer note on that step: "Data will be sent to manual validation." - Otherwise, a "Validate Classification Results" step runs alongside a placeholder comment ("Write your custom Data Classification & Business Rule Validation code here"). The specific rule or confidence threshold applied here is not recoverable from the source — treat it as site-configurable and confirm with the developer before relying on it.
Stage 7 — Human confirmation of document type (only if stage 6 said "not successful")
- Unattended path (
Main-ActionCenter):- Create a document-classification review task in the catalogue named by
ActionCatalog, titledClassificationActionTitle+ the file name, medium priority. - If a queue item exists, then stamp its progress with
TransactionProgress_ClassificationValidation(suffixed with the run'slogKey, so the job processing it can be traced). - Log the task ID, then suspend the job until a reviewer completes the task, and resume with the human-confirmed types. The task activity is set to remove the document from storage on resume and to retry on failure. Developer note flags two production decisions: the file is not downloaded back by default (fill in
TemporaryLocalFolderand the document path arguments if you need it), and removing data from storage may need disabling for auditing, in which case a separate cleanup job is required.
- Create a document-classification review task in the catalogue named by
- Attended path (
Main-Attended): stamp queue progress as above, then present the Classification Station to the operator on screen. - Either way, if the reviewer rejects the document, then a document-rejected exception is raised — see Exceptions.
- If config
SkipClassifierTrainingis false, then retrain the keyword classifier from the confirmed result:- Take a private working copy of the learning file at
ClassifierLearningFilePath+.lock+ the run'slogKey, retrying the copy every second until it exists. - Train against that copy.
- Move the copy back over the original, overwriting, retrying every second until done.
- Both lock and unlock are bounded by
MaxLockTimeoutand set to continue on error. Developer note: "For classifier training, losing a small bit of training data to allow data extraction seems the preferred approach" and "On rare occasions, some training data WILL be lost: when 2 jobs attempt to write at the exact same time, the second job will overwrite the training done by the first."
- Take a private working copy of the learning file at
- Replace the automatic classification results with the human-confirmed ones for the rest of the run.
Stage 8 — Extract fields, once per document found in the file
The remaining stages run in a loop over every classified document (page range) inside the input file. Each pass is individually error-trapped so one bad document does not stop the others.
-
Log the start with the document type and page range.
-
Run the extractors that match the document type against that page range:
Extractor Applies to Notable settings Regex-based extractor Cert of Filing2000 ms timeout, visual alignment off Form extractor W9endpoint FormExtractorEndpoint, API keyApiKey, minimum overlap 65%DU project extractor remaining/configured types 900 000 ms (15 min) timeout, project Predefined, credential asset pathsAcademy/UiDemo_CredentialandShared/ACMECredential -
Retries governed by
MaxExecutionAttemptsExtract; same paid-licence caution. -
Developer note explains a deliberate choice: only the document type is passed to the extraction scope, not the whole classification result, "because the references & token inside the classification result do not match split files."
Stage 9 — Extraction business-rule check
- If config
AlwaysValidateExtractionis true, then force "not successful" and go straight to human validation. - Otherwise, branch on the document type:
Semi-StructuredDocuments.Financial.Invoice→ invoice post-processing (below)Semi-StructuredDocuments.Financial.Receipt→ receipt post-processing (below)- anything else → a placeholder for custom rules; no automated check runs.
- Invoice post-processing. Developer note is emphatic that this is demo-grade: "This should NOT be used as-is, except for demo purposes." It assumes the out-of-the-box invoice field names and EN-US number formatting (
.decimal,,thousands, e.g.10,000.00). In order:- Build a name→value list for every field in the invoice taxonomy, and export the line items to a table. If the mandatory columns are absent, an empty table is deliberately produced so the check below fails.
- Parse the
Datefield against the formats inDU_DateFormats(semicolon-separated). If it is missing or unparseable, then log the failure and mark extraction unsuccessful. - If any field listed in
MandatoryFieldsis empty, or the line-item table has no rows, then log and mark unsuccessful. - Strip currency symbols, spaces and any non-numeric/non-decimal-point characters from the numeric fields so arithmetic is possible.
- For each line item: count the decimal places actually present in
Line Amount, then checkUnit Price × Quantity, rounded to that many places, equalsLine Amount. If it does not, then log (distinguishing "a value was empty" from "the maths does not add up"), mark unsuccessful, and stop checking further lines. Otherwise add the line amount to a running subtotal. - If the running subtotal ≠
Net Amount(tolerated whenNet Amountis absent or zero) then log and mark unsuccessful. - If
Net Amountplus every field named inSubTotalAdditions≠Total, then log and mark unsuccessful. - If any field in
ConfidenceFieldsscores below its own configured threshold, then log the offending field names and mark unsuccessful. - If any remaining field scores below the
other-Confidencethreshold, then log the offending field names and mark unsuccessful. - Only if all of the above pass is extraction marked successful.
- Receipt post-processing. Identical in shape, with three differences: the date field is
Receipt Date, the total field isTotal Value, and there is no per-lineQuantity × Unit Pricecheck — line amounts are simply summed. - If config
AutocorrectionEnabledis true, then run the OCR-mistake correction routine (stage 11) at this point as well, before any human sees the data.
Stage 10 — Human validation of extracted data (only if stage 9 said "not successful")
- Unattended path:
- Log the intent with document type and page range, then create a document-data-validation task in
ActionCatalog, titledValidationActionTitle+ the document type name, medium priority. - If a queue item exists, then stamp its progress with
TransactionProgress_ExtractionValidation+logKey. - Suspend the job until the reviewer completes the task; resume with the corrected values. Same storage-removal and file-download caveats as stage 7.
- Log the intent with document type and page range, then create a document-data-validation task in
- Attended path: stamp progress, then present the Validation Station on screen. Immediately afterwards, copy the document ID from the automatic result over the validated one — developer note: "DocumentId coming from the Validation Station is missing the file extension", and the export file name depends on it.
- If config
AutocorrectionEnabledis true, then run the OCR-mistake correction routine (stage 11), this time with both the before-review and after-review values available so it can learn from what the human changed. - If config
SkipExtractorTrainingis false, then feed the human-validated data to the receipt and invoice ML trainers, writing training output toReceiptsTrainingFolderandInvoicesTrainingFolder. Developer note: the trainers can save locally, to AI Center, or both — local is the default; fill in the Project and Dataset fields to also push to AI Center. This step retries and is set to continue on error: "losing a bit of training data to allow data export seems a preferred approach." - Replace the automatic extraction result with the validated one for export.
Stage 11 — Learn and apply repeat OCR mistakes
This runs only when AutocorrectionEnabled is true, and is invoked from stages 9 and 10. Developer note: "Detect and automatically correct common OCR mistakes for fields with repetitive values (E.g. Logo) by saving corrections made in Action Center. This should NOT be used as-is."
- Read the correction-mapping JSON from storage bucket
StorageBucketName, pathAutocorrectStorageBucketDirectoryPath(retried). - If the file is not valid JSON, then raise an error naming the file — caught by the outer wrapper, which downgrades it to a warning and lets the document continue.
- If this is the pre-review call (no validated result yet), then work from the automatic extraction result alone.
- For each field ID listed in
AutocorrectionFieldIdList(comma-separated):- If the field does not belong to this document type, then log and skip.
- If the field is missing after review, then log and skip.
- If the current value is flagged
Is Black Listedin the mapping, then log and skip. Developer note gives the reason: "There are some generic values that can be read as OCR mistakes that can not be mapped to a single correct value so they should be blacklisted (E.g. 'Bank Of' is an incomplete OCR reading that may be any bank)." - If the current value matches a known mistake, then take the correction with the highest repetition count. If that count has reached
MinimumNumberOfRepetitions, then overwrite the field value (and its confidence) with the correction and log the change; otherwise log that the threshold was not met and leave the value alone. - If a human changed the value during review, then record before→after in the mapping: add a new mistake entry with repetition count 1 if unseen, add a new correction at count 1 if the mistake is known but this correction is not, or increment the existing count.
- Write the updated mapping back to the storage bucket (retried).
- If anything in this stage fails, then log a Warning only (
"Failed to autocorrect ocr mistakes - …"); the document still proceeds.
Stage 12 — Export
- Build the output file name:
ExportsFolder+ the document ID without extension +_+ start page +-+ end page +.xlsx. Developer note: "Adding the page range here ensures a unique export name in case PDF splitting is DISABLED." - Convert the extraction result into one or more tables (confidence and OCR-confidence columns excluded).
- Write each table to its own worksheet in that workbook, named after the table. Wrapped in a retry because the export folder may be a share. Developer note: because this can run inside a parallel loop, plain file writes are used rather than an Excel application session.
- Developer note, important for anyone planning a rebuild: "Please note that the below example is intended for illustrative purposes only. The export to excel was just an example. Ideally you would want to use the data extracted in another process" — with UiPath DataService suggested as the real target.
Stage 13 — Close out
- After every document in the input file has been extracted, validated and exported, mark the queue item Successful.
- A placeholder exists for post-export cleanup; nothing is implemented.
Exceptions and recovery
| Condition | What the automation does | What a human should do |
|---|---|---|
Data\Config.xlsx missing or locked |
No retry. The run fails immediately and without a meaningful message, because the message texts live in that file. | Confirm the file exists at Data\Config.xlsx and nobody has it open exclusively; restart the job. |
An Orchestrator-stored setting on the Assets sheet cannot be fetched |
If a same-named value already came from the spreadsheet, carries on with it. If not, logs an Error naming the asset and aborts the run. | Check the asset exists in the folder named in the OrchestratorAssetFolderPath column and the robot has read rights. Re-run. |
No queue item available, or the item carries no TargetFile |
Logs a warning and ends cleanly. No document processed, no failure status recorded. | Nothing, if the queue is legitimately empty. If items exist but lack TargetFile, take it up with whoever loads the queue — that loader is outside this automation. |
| OCR, classifier or extractor call fails transiently | Each stage retries independently using MaxExecutionAttemptsDigitize / MaxExecutionAttemptsClassify / MaxExecutionAttemptsExtract and the configured interval. |
Watch the licence cost. Developer flags that retrying a paid OCR/classifier/extractor call may be billed again; do not raise the retry counts casually. |
Invoice/receipt date missing or unparseable against DU_DateFormats |
Logs the failure, marks extraction unsuccessful, routes the document to human validation. | Reviewer enters the correct date. If it recurs for a given supplier's format, add that format to DU_DateFormats. |
| Mandatory field or mandatory table column missing, or no line items found | Logs, marks unsuccessful, routes to human validation. | Reviewer fills the gaps. Persistent misses point at extractor tuning or a wrong MandatoryFields / MandatoryColumns list. |
| Line arithmetic fails, subtotal ≠ Net Amount, or Net Amount + additions ≠ Total | Logs (distinguishing empty values from a genuine mismatch), stops checking further lines, routes to human validation. | Reviewer corrects the figures. Check whether SubTotalAdditions is missing a tax or freight field for this document layout. |
Any field below its own confidence threshold, or any other field below other-Confidence |
Logs the failing field names, routes to human validation. | Reviewer confirms or corrects. Repeat offenders are candidates for extractor retraining. |
| A reviewer rejects the document at classification or data validation | A document-rejected exception is raised, caught by the per-document handler, and logged as a Warning with page range, file name, exception type and message. Processing continues with the remaining documents in the same file. Developer note: "Exceptions should NEVER be rethrown here! A rethrown exception would stop the processing flow of ALL documents within the input file and would orphan any pending" reviews. | Investigate why it was rejected — wrong document, unreadable scan, out of scope. The rejected document produces no export; re-submit it once resolved. |
| Correction-mapping file in the storage bucket is not valid JSON | Raises an error internally, but the wrapper downgrades it to a Warning; the document is processed normally without autocorrection. | Repair or replace the JSON at AutocorrectStorageBucketDirectoryPath in bucket StorageBucketName. Until then, learned corrections are silently not being applied or saved. |
| Two robots update the classifier learning file at the same time | Each copies the file to a per-run lock name (ClassifierLearningFilePath + .lock + logKey), trains against the copy, then moves it back with overwrite. Bounded by MaxLockTimeout, continue-on-error. Developer admits: "On rare occasions, some training data WILL be lost." |
Accept the risk, or switch to the supplied queue-based single-token alternative (GetWritePermission / GiveUpWritePermission), which needs a dedicated AccessQueue holding exactly one item with automatic retries disabled. Nothing in the current main flow calls it. |
| Setting the queue item's progress marker fails | Retried, then ignored (continue-on-error) — "failing to update the transaction progress should not stop the processing of the transaction." | Nothing. Progress markers are diagnostic only. |
| Any unhandled error at run level | Logs an Error with exception type, message and stack trace; runs cleanup; then sets the queue item Failed — error type Business for a business-rule or document-rejected exception (deliberately not retried, since "the result will be the same until the problem that causes the exception is solved"), Application for anything else (retryable after restarting the applications involved); then rethrows so the job itself shows as faulted. | Read the log line carrying the run's logKey. Business failures need the underlying data or rule fixed before retrying. Application failures can usually be retried from the queue after restarting the affected systems. |
Data handled
| Name | What it holds | Where it comes from |
|---|---|---|
Settings list (config) |
Every runtime setting, message text and threshold, as name→value pairs. Also acts as the run's scratchpad — queue-item fields and Orchestrator asset values are merged into it. | Data\Config.xlsx sheets Settings, Constants, AutocorrectOcrMistakes, InvoicePostProcessing, ReceiptPostProcessing; then overwritten by Assets-sheet values from Orchestrator; then overwritten again by fields on the queue item |
logKey |
Unique tag for the run, appended to every log line and every queue-item progress note | Generated at config load |
maxAttempts, retryInterval |
Default retry behaviour for Orchestrator and file operations. Never below 1 attempt. | Derived from config |
TransactionItem |
The queue item being worked, used later to set progress and final status | Orchestrator queue DocumentUnderstandingQueueName |
in_TargetFile |
Full path of the document being processed | Queue item's TargetFileKey field, or the in_TargetFile argument, or the attended file picker |
docTaxonomy |
The catalogue of document types and their field names — drives which fields the post-processing checks look for | Loaded at initialisation (stage 2) |
dom, docText |
The OCR output: page/word layout model and plain text. Passed to classification, extraction and both trainers. | Stage 4 digitisation |
classificationResultsArray |
One entry per document found in the file: document type plus page range. Replaced wholesale by the human-confirmed version if a reviewer intervenes. | Intelligent Keyword Classifier (stage 5); or validatedClassificationResults after stage 7 |
classificationSuccessFlag |
Whether classification can be trusted without a human | Stage 6 rule check |
currentPageRange |
Human-readable page span (e.g. 3-5) of the document currently being handled; used in logs, review-task context and export file names |
Computed from the classification result's bounds |
extractionResults |
All extracted field values and line-item tables for one document. Replaced by the validated version after human review. | Stage 8 extractors; or validatedExtractionResults after stage 10 |
extractionSuccessFlag |
Whether the extraction passed every arithmetic and confidence rule | Stage 9 post-processing |
documentFields |
Field name → extracted raw value for the document under check; the basis of every post-processing rule | Built by walking the taxonomy against the extraction result |
itemTable |
The line items as a table. Deliberately produced empty if a mandatory column is absent, so the mandatory-fields check fails. | Extraction result exported to a dataset |
mandatoryFields, mandatoryColumnFields, subTotalAdditions, confidenceFields, otherConfidenceFields |
The rule lists: which fields must be present, which columns the line-item table must have, which amounts add to Net Amount to reach Total, which fields have their own confidence threshold, and which fall back to other-Confidence |
InvoicePostProcessing / ReceiptPostProcessing sheets of Config.xlsx |
subtotal, total |
Running sums used to cross-check Net Amount and Total |
Computed from the line-item table |
correctionDataJson |
The learned OCR-mistake mapping: per field, each misread value, its candidate corrections, each correction's repetition count, and a blacklist flag | Storage bucket StorageBucketName, path AutocorrectStorageBucketDirectoryPath; written back after each run |
valueBeforeHitL, valueAfterHitL |
A field's value before and after human review — the raw material for learning corrections | Automatic vs validated extraction results |
keywordLockName |
Path of the private working copy of the classifier learning file: ClassifierLearningFilePath + .lock + logKey |
Built in LockFile.xaml |
outputPath |
Destination workbook: ExportsFolder + document ID + _ + page range + .xlsx |
Built in stage 12 |
Not determinable from the source
- Who owns and fills the queue. No dispatcher or loader workflow exists here; the
TargetFilepaths arrive from something outside this automation. - Schedule and volume. No trigger definition, and each run handles exactly one item — whether it is looped, run per job, or timed is not visible.
- Actual config values.
Config.xlsxis not in the source. Unknown concretely: the queue name and folder path,ExportsFolder,ActionCatalog,StorageBucketName,AutocorrectStorageBucketDirectoryPath,ClassifierLearningFilePath,InvoicesTrainingFolder,ReceiptsTrainingFolder,MandatoryFields,MandatoryColumns,SubTotalAdditions,ConfidenceFieldsand their thresholds,other-Confidence,DU_DateFormats,MinimumNumberOfRepetitions,AutocorrectionFieldIdList,MaxLockTimeout, all retry counts, and the on/off flagsAlwaysValidateClassification,AlwaysValidateExtraction,AutocorrectionEnabled,SkipClassifierTraining,SkipExtractorTraining. - The automated classification rule. Stage 6 contains a "Validate Classification Results" step plus a placeholder comment, so the confidence threshold or logic that decides classification is trustworthy cannot be recovered.
- What consumes the export. The developer explicitly calls the Excel export illustrative and recommends a data service instead; the real downstream consumer is unstated.
- Who the reviewers are, their turnaround expectations, and how tasks are routed to them — the task titles come from config values not present in the source.
- Which document types are genuinely in scope. Extractors exist for Cert of Filing, W9, invoices and receipts, and the DU project extractor lists projects (
Predefined,Test_invoice,Delivery Notes,Invoice,LALA,TestProject,Test_BestClassification,AL_Test1) that look like development leftovers. - Which entry point is deployed — unattended
Main-ActionCenter.xamlor attendedMain-Attended.xaml. Both exist and are otherwise near-identical. - Retention and audit policy for reviewed documents. Both review activities remove data from storage on resume; the developer notes this may need disabling in production with a separate cleanup, but the decision is not recorded.
- Whether the queue-based single-token locking alternative is in use (
GetWritePermission/GiveUpWritePermissionand its prerequisiteAccessQueue), or only file-copy locking — nothing in the main flow calls it. - How many robots run concurrently, which determines whether the acknowledged risk of losing classifier training data is material.
How this SOP was checked
Generated from the project's source files and audited against them. Audit verdict: minor issues, confidence high.
Faithful end-to-end trace of the real business flow: config load (sheets + Assets overlay), taxonomy/asset init with retry, queue item retrieval and TargetFileKey handling, digitisation, keyword classification with splitting, both business-rule gates, both human-review paths (Action Center suspend vs attended stations), classifier/extractor training with file-copy locking, autocorrect-OCR learning loop, Excel export naming, EndProcess status, and the Business-vs-Application status logic in SetTransactionStatus. Config values, owner, schedule, reviewer routing are correctly declared unrecoverable. Not covered (and arguably out of scope): the Tests/* regression suite and Mocks/* variants (~25 assets) that compare digitisation/classification/extraction against cached DOM, text and results, plus the batch caching utilities — an analyst rebuilding the bot would want these mentioned.
Remaining minor notes:
- Stage 7 step 3 and Exceptions table row 'A reviewer rejects the document at classification or data validation' — Classification-stage rejection is attributed to the per-document handler (ERR_HandleDocumentError, Warn, processing continues). In both Mains the classification validation block sits outside the 'Process Each Document' loop, inside the outer 'Try Catch - Process File'. A DocumentRejectedByUserException there goes to ERR_AbortProcess (Error log), SetTransactionStatus with Failed/Business, and is rethrown — the whole run aborts. (Split the row: data-validation rejection inside the per-document loop is warned and skipped (other documents continue); classification rejection aborts the run, marks the queue item Failed/Business and faults the job.)
- Stage 9 step 4 (receipt post-processing) — Described as 'identical in shape' to the invoice rules apart from three differences, implying the subtotal-vs-
Net Amountcheck (step 3.6) also applies. ReceiptPostProcessing's Extraction Results Check contains only Total Issue? → specific-field confidence → other-field confidence; there is no Net Amount comparison. Also, in the source the receipt 'Total Calculation' (subtotal + SubTotalAdditions) is executed in the False branch after the checks, so the Total comparison runs against an un-summed value. (State that receipts have no subtotal-vs-Net-Amount check, and note the odd placement of the total calculation after the comparison in the source flowchart.) - Stage 1 steps 3–4 — Ordering: the outline initialises the config dictionary and logKey ('Assign - Initialize Config and LogKey') before the sheet loop, and derives MaxExecutionAttempts/RetryInterval last. The SOP derives retry values (step 3) then generates logKey (step 4). (Order as: log start → initialise config and logKey → read the five sheets → derive MaxAttempts/RetryInterval (floor 1).)
- Stage 8 extractor table ('Applies to' column) — The mapping of the regex extractor to 'Cert of Filing', the form extractor to 'W9', and the DU project extractor to 'remaining/configured types' is inferred from extractor display names; the outline shows no taxonomy-to-extractor mapping. (Present the extractor names as-is and flag that which document type each extractor is bound to is not recoverable from the source.)
- Stage 6 / Stage 9 step ordering within the business-rule workflows — Both 35_ and 55_ are flowcharts whose start node is not visible in the outline; the 'LogMessage_...BusinessRuleValidationStart' log node and (in 55) the autocorrect node are chained separately from the AlwaysValidate branch. The SOP presents a definite order (branch first, autocorrect last) without flagging the ambiguity, and omits the start log entirely in Stage 6. (Mention the start-of-validation log entry and note that the exact node order inside these two flowcharts is not determinable from the source.)