Upload PDF bank statements
Upload up to 10 bank statements for a customer in one job and get back every transaction, extracted and classified.
Before you start
- An API key with the
transactions:extractscope (see Authentication). The Node.js, Python and C# snippets read it from theWSD_API_KEYenvironment variable. - The sandbox base URL:
https://sandbox.walkerstdata.com.au. - A
customerId(see Classify a customer's transactions). - Machine-readable statements from a supported bank, not scans or photos.
- Node.js snippets use the built-in
fetch(Node 18+) and top-levelawait, so save them as an.mjsfile. Python snippets userequests. C# snippets are .NET 8 top-level programs usingHttpClientandSystem.Text.Json, with no extra packages.
POST /v1/customer/{customerId}/transactions/pdf takes a single PDF as multipart/form-data (field name file) and returns a jobId in one call:
- cURL
- Node.js
- Python
- C#
curl -X POST "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions/pdf" \
-H "x-api-key: YOUR_API_KEY" \
-F "file=@cba-business-statement-2025-04.pdf"
import { readFile } from 'node:fs/promises';
const BASE_URL = 'https://sandbox.walkerstdata.com.au';
const HEADERS = { 'x-api-key': process.env.WSD_API_KEY };
const customerId = '524ba9bc-06e7-417e-bd54-b99785f5194a';
const path = 'cba-business-statement-2025-04.pdf';
const form = new FormData();
const pdf = new Blob([await readFile(path)], { type: 'application/pdf' });
form.append('file', pdf, path);
const resp = await fetch(
`${BASE_URL}/v1/customer/${customerId}/transactions/pdf`,
{ method: 'POST', headers: HEADERS, body: form }
);
if (!resp.ok) throw new Error(`Upload failed: ${resp.status}`);
const jobId = (await resp.json()).data.jobId;
import os
import time
from pathlib import Path
import requests
BASE_URL = "https://sandbox.walkerstdata.com.au"
HEADERS = {"x-api-key": os.environ["WSD_API_KEY"]}
customer_id = "524ba9bc-06e7-417e-bd54-b99785f5194a"
path = Path("cba-business-statement-2025-04.pdf")
resp = requests.post(
f"{BASE_URL}/v1/customer/{customer_id}/transactions/pdf",
headers=HEADERS,
files={"file": (path.name, path.read_bytes(), "application/pdf")},
)
resp.raise_for_status()
job_id = resp.json()["data"]["jobId"]
using System.Net.Http.Headers;
using System.Net.Http.Json;
using System.Text.Json.Nodes;
var http = new HttpClient { BaseAddress = new Uri("https://sandbox.walkerstdata.com.au") };
http.DefaultRequestHeaders.Add("x-api-key", Environment.GetEnvironmentVariable("WSD_API_KEY"));
var customerId = "524ba9bc-06e7-417e-bd54-b99785f5194a";
var path = "cba-business-statement-2025-04.pdf";
var pdf = new ByteArrayContent(await File.ReadAllBytesAsync(path));
pdf.Headers.ContentType = new MediaTypeHeaderValue("application/pdf");
using var form = new MultipartFormDataContent { { pdf, "file", Path.GetFileName(path) } };
var resp = await http.PostAsync($"/v1/customer/{customerId}/transactions/pdf", form);
resp.EnsureSuccessStatusCode();
var jobId = (string)(await resp.Content.ReadFromJsonAsync<JsonNode>())!["data"]!["jobId"]!;
Then skip to step 4. The multi-file flow below suits several statements, or files you'd rather send straight to storage.
1. Create the upload job
List the files you're going to upload (1 to 10, each ending in .pdf, .html or .htm). The response gives you a job, an uploadUrl and a blobName for each file, in the same order as your request.
- cURL
- Node.js
- Python
- C#
curl -X POST "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions/multi-pdf" \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"files": [
{ "fileName": "cba-business-statement-2025-04.pdf" },
{ "fileName": "cba-business-statement-2025-05.pdf" }
]
}'
import { readFile } from 'node:fs/promises';
import { basename } from 'node:path';
const BASE_URL = 'https://sandbox.walkerstdata.com.au';
const HEADERS = {
'x-api-key': process.env.WSD_API_KEY,
'Content-Type': 'application/json',
};
const customerId = '524ba9bc-06e7-417e-bd54-b99785f5194a';
const paths = [
'cba-business-statement-2025-04.pdf',
'cba-business-statement-2025-05.pdf',
];
const resp = await fetch(
`${BASE_URL}/v1/customer/${customerId}/transactions/multi-pdf`,
{
method: 'POST',
headers: HEADERS,
body: JSON.stringify({
files: paths.map((p) => ({ fileName: basename(p) })),
}),
}
);
if (!resp.ok) throw new Error(`Create upload failed: ${resp.status}`);
const upload = (await resp.json()).data;
const jobId = upload.jobId;
import os
import time
from pathlib import Path
from urllib.parse import quote, urlsplit, urlunsplit
import requests
BASE_URL = "https://sandbox.walkerstdata.com.au"
HEADERS = {"x-api-key": os.environ["WSD_API_KEY"]}
customer_id = "524ba9bc-06e7-417e-bd54-b99785f5194a"
paths = [Path("cba-business-statement-2025-04.pdf"), Path("cba-business-statement-2025-05.pdf")]
resp = requests.post(
f"{BASE_URL}/v1/customer/{customer_id}/transactions/multi-pdf",
headers=HEADERS,
json={"files": [{"fileName": p.name} for p in paths]},
)
resp.raise_for_status()
upload = resp.json()["data"]
job_id = upload["jobId"]
using System.Net.Http.Headers;
using System.Net.Http.Json;
using System.Text.Json.Nodes;
var http = new HttpClient { BaseAddress = new Uri("https://sandbox.walkerstdata.com.au") };
http.DefaultRequestHeaders.Add("x-api-key", Environment.GetEnvironmentVariable("WSD_API_KEY"));
var customerId = "524ba9bc-06e7-417e-bd54-b99785f5194a";
string[] paths = ["cba-business-statement-2025-04.pdf", "cba-business-statement-2025-05.pdf"];
var resp = await http.PostAsJsonAsync(
$"/v1/customer/{customerId}/transactions/multi-pdf",
new { files = paths.Select(p => new { fileName = Path.GetFileName(p) }) });
resp.EnsureSuccessStatusCode();
var upload = (await resp.Content.ReadFromJsonAsync<JsonNode>())!["data"]!;
var jobId = (string)upload["jobId"]!;
{
"data": {
"jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863",
"uploadUrl": "https://examplestorage.blob.core.windows.net/bank-statements?sv=2024-08-04&se=2026-09-25T01%3A40%3A00Z&sr=c&sp=cw&sig=EXAMPLEsignature%3D",
"files": [
{
"taskId": "b7e2c4d9-5a1f-4e38-9c6b-2d8f0a3e71c5",
"blobName": "statements/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/cba-business-statement-2025-04-00.pdf"
},
{
"taskId": "19f6b2d8-c3a7-4e50-8d4b-a2e7f0c61935",
"blobName": "statements/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/cba-business-statement-2025-05-01.pdf"
}
]
},
"message": null
}
2. Upload each file
uploadUrl is a write-only Azure Blob Storage SAS URL for the statements container, valid for 20 minutes. For each file, build its URL by adding /<blobName> to the path of uploadUrl (keeping the query string), then PUT the file's bytes there with the x-ms-blob-type: BlockBlob header. Don't send your x-api-key to storage: the SAS query string is the credential. Azure responds 201 Created for each successful upload.
- cURL
- Node.js
- Python
- C#
UPLOAD_URL='https://examplestorage.blob.core.windows.net/bank-statements?sv=2024-08-04&se=...&sig=...'
BLOB_NAME='statements/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/cba-business-statement-2025-04-00.pdf'
curl -X PUT "${UPLOAD_URL%%\?*}/${BLOB_NAME}?${UPLOAD_URL#*\?}" \
-H "x-ms-blob-type: BlockBlob" \
-H "Content-Type: application/pdf" \
--data-binary @cba-business-statement-2025-04.pdf
Repeat for each file with its own blobName.
for (const [i, target] of upload.files.entries()) {
const blobUrl = new URL(upload.uploadUrl);
const encoded = target.blobName.split('/').map(encodeURIComponent).join('/');
blobUrl.pathname = `${blobUrl.pathname}/${encoded}`;
const put = await fetch(blobUrl, {
method: 'PUT',
headers: {
'x-ms-blob-type': 'BlockBlob',
'Content-Type': 'application/pdf',
},
body: await readFile(paths[i]),
});
if (!put.ok) throw new Error(`Upload of ${paths[i]} failed: ${put.status}`);
}
container = urlsplit(upload["uploadUrl"])
for path, target in zip(paths, upload["files"]):
blob_url = urlunsplit(
(container.scheme, container.netloc,
f"{container.path}/{quote(target['blobName'])}", container.query, "")
)
put = requests.put(
blob_url,
data=path.read_bytes(),
headers={"x-ms-blob-type": "BlockBlob", "Content-Type": "application/pdf"},
)
put.raise_for_status()
using var storage = new HttpClient(); // no x-api-key: the SAS query string is the credential
var container = new Uri((string)upload["uploadUrl"]!);
var targets = upload["files"]!.AsArray();
for (var i = 0; i < paths.Length; i++)
{
var blobName = (string)targets[i]!["blobName"]!;
var encoded = string.Join("/", blobName.Split('/').Select(Uri.EscapeDataString));
var blobUrl = new Uri($"{container.GetLeftPart(UriPartial.Path)}/{encoded}{container.Query}");
var content = new ByteArrayContent(await File.ReadAllBytesAsync(paths[i]));
content.Headers.ContentType = new MediaTypeHeaderValue("application/pdf");
using var put = new HttpRequestMessage(HttpMethod.Put, blobUrl) { Content = content };
put.Headers.Add("x-ms-blob-type", "BlockBlob");
var uploaded = await storage.SendAsync(put);
if (!uploaded.IsSuccessStatusCode)
{
throw new Exception($"Upload of {paths[i]} failed: {(int)uploaded.StatusCode}");
}
}
If the URL expires before you finish, start again from step 1 with a new upload job.
3. Seal the upload
Once every file is uploaded, seal the job. Walker Street Data checks that all the expected files are there, then starts extraction. Sealing is idempotent, so it's safe to retry.
- cURL
- Node.js
- Python
- C#
curl -X POST "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions/multi-pdf/seal" \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863" }'
const seal = await fetch(
`${BASE_URL}/v1/customer/${customerId}/transactions/multi-pdf/seal`,
{ method: 'POST', headers: HEADERS, body: JSON.stringify({ jobId }) }
);
if (!seal.ok) throw new Error(`Seal failed: ${seal.status}`);
resp = requests.post(
f"{BASE_URL}/v1/customer/{customer_id}/transactions/multi-pdf/seal",
headers=HEADERS,
json={"jobId": job_id},
)
resp.raise_for_status()
var seal = await http.PostAsJsonAsync(
$"/v1/customer/{customerId}/transactions/multi-pdf/seal", new { jobId });
seal.EnsureSuccessStatusCode();
{
"data": {
"jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863",
"status": "Processing",
"totalTasks": 6,
"extractionTasks": 2,
"enrichmentTasks": 1
},
"message": null
}
If a file is missing, the seal returns 400 and lists the files still to upload in data.missingFiles (each with its taskId and blobName). Upload those and seal again.
4. Wait for the job to finish
Poll the job until its status is Completed, CompletedWithErrors or Failed, or use a webhook instead. Extraction takes longer than a JSON submission, so poll every 10 seconds or so.
- cURL
- Node.js
- Python
- C#
curl "https://sandbox.walkerstdata.com.au/v1/jobs/3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863/status" \
-H "x-api-key: YOUR_API_KEY"
const TERMINAL = ['Completed', 'CompletedWithErrors', 'Failed'];
let job;
while (true) {
const r = await fetch(`${BASE_URL}/v1/jobs/${jobId}/status`, {
headers: HEADERS,
});
if (!r.ok) throw new Error(`Job status failed: ${r.status}`);
job = (await r.json()).data;
if (TERMINAL.includes(job.status)) break;
await new Promise((resolve) => setTimeout(resolve, 10000));
}
TERMINAL = {"Completed", "CompletedWithErrors", "Failed"}
while True:
resp = requests.get(f"{BASE_URL}/v1/jobs/{job_id}/status", headers=HEADERS)
resp.raise_for_status()
job = resp.json()["data"]
if job["status"] in TERMINAL:
break
time.sleep(10)
string[] terminal = ["Completed", "CompletedWithErrors", "Failed"];
JsonNode job;
while (true)
{
job = (await http.GetFromJsonAsync<JsonNode>($"/v1/jobs/{jobId}/status"))!["data"]!;
if (terminal.Contains((string?)job["status"])) break;
await Task.Delay(TimeSpan.FromSeconds(10));
}
5. Check for extraction problems
A statement can be partly readable, or not readable at all, and the upload still succeeds; problems only show up on the finished job. Before you rely on the data, check:
failedDocuments: files that couldn't be processed, each with itsfileNameanderrorwarningsandfileWarnings: issues such as unreconciled rows or a running balance that doesn't match the statement (see Extraction warnings)
{
"data": {
"jobId": "3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863",
"status": "CompletedWithErrors",
"flow": "PdfUpload",
"totalReceivedTransactions": 412,
"totalProcessedTransactions": 398,
"totalDuplicateTransactions": 14,
"warnings": {
"warnings": [],
"rowDiscrepancies": ["Row 44: amount could not be parsed"],
"missingAccountDetails": [],
"other": [],
"balanceChainMismatch": {
"pdf": 23725.33,
"computed": 23675.33,
"delta": -50.0
}
},
"fileWarnings": [
{
"fileName": "cba-business-statement-2025-04-00.pdf",
"warnings": ["Row 44: amount could not be parsed"]
}
],
"failedDocuments": null
},
"message": null
}
6. Read the transactions
Fetch the extracted, classified transactions for this job. The response has the same shape as for JSON submissions: see Classify a customer's transactions for an example.
- cURL
- Node.js
- Python
- C#
curl "https://sandbox.walkerstdata.com.au/v1/customer/524ba9bc-06e7-417e-bd54-b99785f5194a/transactions?jobId=3f9a6c2e-81d4-4b7a-a5e0-c7d219f4b863&page=1&count=100" \
-H "x-api-key: YOUR_API_KEY"
const params = new URLSearchParams({ jobId, page: '1', count: '100' });
const txns = await fetch(
`${BASE_URL}/v1/customer/${customerId}/transactions?${params}`,
{ headers: HEADERS }
);
if (!txns.ok) throw new Error(`Get transactions failed: ${txns.status}`);
const { data } = await txns.json();
console.log(
data.totalCount,
'transactions across',
data.accounts.length,
'account(s)'
);
resp = requests.get(
f"{BASE_URL}/v1/customer/{customer_id}/transactions",
headers=HEADERS,
params={"jobId": job_id, "page": 1, "count": 100},
)
resp.raise_for_status()
data = resp.json()["data"]
print(data["totalCount"], "transactions across", len(data["accounts"]), "account(s)")
var data = (await http.GetFromJsonAsync<JsonNode>(
$"/v1/customer/{customerId}/transactions?jobId={jobId}&page=1&count=100"))!["data"]!;
Console.WriteLine($"{data["totalCount"]} transactions across {data["accounts"]!.AsArray().Count} account(s)");
Accounts found on the statements appear in accounts with sourcedVia set to pdf.
What's next
- Upload a PDF bank statement: supported banks and how failures surface
- Get notified when a job finishes: skip the polling loop
- Get a customer's reports and a lending decision: use the extracted data
- API reference: Submit multi-PDF, Seal multi-PDF upload, Submit PDF, Get job status