Skip to main content

Upload PDF bank statements

Upload up to 10 bank statements for a customer in one job and get back every transaction, extracted and classified.

Before you start​

  • An API key with the transactions:extract scope (see Authentication). The Node.js, Python and C# snippets read it from the WSD_API_KEY environment variable.
  • The sandbox base URL: https://sandbox.walkerstdata.com.au.
  • A customerId (see Classify a customer's transactions).
  • Machine-readable statements from a supported bank, not scans or photos.
  • Node.js snippets use the built-in fetch (Node 18+) and top-level await, so save them as an .mjs file. Python snippets use requests. C# snippets are .NET 8 top-level programs using HttpClient and System.Text.Json, with no extra packages.
Only one statement?

POST /v1/customer/{customerId}/transactions/pdf takes a single PDF as multipart/form-data (field name file) and returns a jobId in one call:

curl -X POST "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions/pdf" \
-H "x-api-key: YOUR_API_KEY" \
-F "file=@cba-business-statement-2025-04.pdf"

Then skip to step 4. The multi-file flow below suits several statements, or files you'd rather send straight to storage.

1. Create the upload job​

List the files you're going to upload (1 to 10, each ending in .pdf, .html or .htm). The response gives you a job, an uploadUrl and a blobName for each file, in the same order as your request.

curl -X POST "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions/multi-pdf" \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"files": [
{ "fileName": "cba-business-statement-2025-04.pdf" },
{ "fileName": "cba-business-statement-2025-05.pdf" }
]
}'
{
"data": {
"jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863",
"uploadUrl": "https://examplestorage.blob.core.windows.net/bank-statements?sv=2024-08-04&se=2026-09-25T01%3A40%3A00Z&sr=c&sp=cw&sig=EXAMPLEsignature%3D",
"files": [
{
"taskId": "b7e2c4d9-5a1f-4e38-9c6b-2d8f0a3e71c5",
"blobName": "statements/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/cba-business-statement-2025-04-00.pdf"
},
{
"taskId": "19f6b2d8-c3a7-4e50-8d4b-a2e7f0c61935",
"blobName": "statements/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/cba-business-statement-2025-05-01.pdf"
}
]
},
"message": null
}

2. Upload each file​

uploadUrl is a write-only Azure Blob Storage SAS URL for the statements container, valid for 20 minutes. For each file, build its URL by adding /<blobName> to the path of uploadUrl (keeping the query string), then PUT the file's bytes there with the x-ms-blob-type: BlockBlob header. Don't send your x-api-key to storage: the SAS query string is the credential. Azure responds 201 Created for each successful upload.

UPLOAD_URL='https://examplestorage.blob.core.windows.net/bank-statements?sv=2024-08-04&se=...&sig=...'
BLOB_NAME='statements/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/cba-business-statement-2025-04-00.pdf'

curl -X PUT "${UPLOAD_URL%%\?*}/${BLOB_NAME}?${UPLOAD_URL#*\?}" \
-H "x-ms-blob-type: BlockBlob" \
-H "Content-Type: application/pdf" \
--data-binary @cba-business-statement-2025-04.pdf

Repeat for each file with its own blobName.

If the URL expires before you finish, start again from step 1 with a new upload job.

3. Seal the upload​

Once every file is uploaded, seal the job. Walker Street Data checks that all the expected files are there, then starts extraction. Sealing is idempotent, so it's safe to retry.

curl -X POST "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions/multi-pdf/seal" \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863" }'
{
"data": {
"jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863",
"status": "Processing",
"totalTasks": 6,
"extractionTasks": 2,
"enrichmentTasks": 1
},
"message": null
}

If a file is missing, the seal returns 400 and lists the files still to upload in data.missingFiles (each with its taskId and blobName). Upload those and seal again.

4. Wait for the job to finish​

Poll the job until its status is Completed, CompletedWithErrors or Failed, or use a webhook instead. Extraction takes longer than a JSON submission, so poll every 10 seconds or so.

curl "https://sandbox.walkerstdata.com.au/v1/jobs/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/status" \
-H "x-api-key: YOUR_API_KEY"

5. Check for extraction problems​

A statement can be partly readable, or not readable at all, and the upload still succeeds; problems only show up on the finished job. Before you rely on the data, check:

  • failedDocuments: files that couldn't be processed, each with its fileName and error
  • warnings and fileWarnings: issues such as unreconciled rows or a running balance that doesn't match the statement (see Extraction warnings)
{
"data": {
"jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863",
"status": "CompletedWithErrors",
"flow": "PdfUpload",
"totalReceivedTransactions": 412,
"totalProcessedTransactions": 398,
"totalDuplicateTransactions": 14,
"warnings": {
"warnings": [],
"rowDiscrepancies": ["Row 44: amount could not be parsed"],
"missingAccountDetails": [],
"other": [],
"balanceChainMismatch": {
"pdf": 23725.33,
"computed": 23675.33,
"delta": -50.0
}
},
"fileWarnings": [
{
"fileName": "cba-business-statement-2025-04-00.pdf",
"warnings": ["Row 44: amount could not be parsed"]
}
],
"failedDocuments": null
},
"message": null
}

6. Read the transactions​

Fetch the extracted, classified transactions for this job. The response has the same shape as for JSON submissions: see Classify a customer's transactions for an example.

curl "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions?jobId=3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863&page=1&count=100" \
-H "x-api-key: YOUR_API_KEY"

Accounts found on the statements appear in accounts with sourcedVia set to pdf.

What's next​