Extract Text and Data from a PDF with a Transformation
When a Source receives an application/pdf request, a Transformation receives the PDF as a Buffer in request.body. You can parse it with an npm library and replace the body with JSON, so the Event becomes searchable and filterable and your Destination receives structured data instead of a file. The parsing runs in Hookdeck, so you don't need to run a document parser in your own service.
This guide extracts an invoice number and total from the text of a PDF with unpdf, and reads document metadata such as the title, author and page count with pdf-lib.
Prerequisites
- A Connection whose source receives PDF files with
Content-Type: application/pdf - Node.js and npm on your machine, to bundle the transformation
Transformations can only import Node.js built-in modules, so you bundle the library into your code with esbuild. See bundle dependencies for background.
Extract text and invoice fields
Install the dependencies
mkdir pdf-transformation && cd pdf-transformation
npm init -y
npm install unpdf
npm install --save-dev esbuild
Write the transformation
Save this as pdf-text.js. It extracts the text of every page, then reads the invoice number and total with regular expressions. Adjust the expressions to the layout of your documents.
import { extractText, getDocumentProxy } from 'unpdf';
export default {
async transform(request) {
const pdf = await getDocumentProxy(new Uint8Array(request.body));
const { totalPages, text } = await extractText(pdf, { mergePages: true });
const invoiceNumber = text.match(/Invoice number:\s*(\S+)/)?.[1] ?? null;
const total = text.match(/Total due:\s*\$([\d,]+\.\d{2})/)?.[1];
request.body = {
invoice_number: invoiceNumber,
total: total ? Number(total.replace(/,/g, '')) : null,
pages: totalPages,
text,
};
request.headers['content-type'] = 'application/json';
return request;
},
};
new Uint8Array(request.body) copies the bytes into a standalone Uint8Array, the type unpdf expects. A Buffer can be a view into a larger shared ArrayBuffer, and the copy keeps the parser from seeing anything outside the file.
Bundle it
npx esbuild pdf-text.js --bundle --format=esm --platform=node --minify --outfile=dist/pdf-text.js
The bundle includes the PDF parser and is about 1.6 MB minified, within the 5 MB code limit.
Add it to your connection
Paste the contents of dist/pdf-text.js into a new transformation on your connection. See create a transformation. You can also upload the file with the Hookdeck CLI:
hookdeck gateway transformation upsert pdf-text --code-file dist/pdf-text.js
Send a PDF
curl -X POST "https://hkdk.events/src_xxxxxxxx" \
-H "Content-Type: application/pdf" \
--data-binary @invoice.pdf
Use --data-binary, not -d: -d strips newlines and corrupts binary files.
What your destination receives
For a two-page invoice, your destination receives Content-Type: application/json and a body like this one:
{
"invoice_number": "INV-1042",
"total": 1250,
"pages": 2,
"text": "ACME Corp\nInvoice number: INV-1042\nInvoice date: 2026-10-01\nBill to: Jane Doe\nWidget x 2 $500.00\nSupport plan $250.00\nTotal due: $1,250.00\nPayment terms: Net 30"
}
Because the event is JSON, you can filter on total or search for invoice_number in the dashboard.
Read document metadata
To read metadata without extracting text, use pdf-lib. It's smaller than unpdf and faster for large documents.
npm install pdf-lib
import { PDFDocument } from 'pdf-lib';
export default {
async transform(request) {
const pdf = await PDFDocument.load(request.body, { updateMetadata: false });
request.body = {
title: pdf.getTitle() ?? null,
author: pdf.getAuthor() ?? null,
page_count: pdf.getPageCount(),
created_at: pdf.getCreationDate()?.toISOString() ?? null,
size: request.body.length,
};
request.headers['content-type'] = 'application/json';
return request;
},
};
npx esbuild pdf-metadata.js --bundle --format=esm --platform=node --minify --outfile=dist/pdf-metadata.js
For the same invoice, your destination receives:
{
"title": "Invoice INV-1042",
"author": "Acme Billing",
"page_count": 2,
"created_at": "2026-10-08T21:21:22.000Z",
"size": 1377
}
updateMetadata: false keeps pdf-lib from changing the producer and modification date as it loads the document. Fields the PDF doesn't set come back as null.
Keep the original PDF
Replacing the body with JSON means your destination no longer receives the PDF. If it needs both, include the file in the JSON as base64:
import { extractText, getDocumentProxy } from 'unpdf';
export default {
async transform(request) {
const original = request.body;
const pdf = await getDocumentProxy(new Uint8Array(original));
const { text } = await extractText(pdf, { mergePages: true });
request.body = {
invoice_number: text.match(/Invoice number:\s*(\S+)/)?.[1] ?? null,
pdf_base64: original.toString('base64'),
};
request.headers['content-type'] = 'application/json';
return request;
},
};
Base64 makes the file about a third larger. Other options:
- Add a second connection on the same source without the transformation, so one destination receives the PDF and another receives the JSON.
- Leave
request.bodyunchanged and put the extracted fields in headers, such asrequest.headers['x-invoice-number']. Returning the body unchanged keeps the original bytes.
Limits and pitfalls
- Execution time. Transformations must finish within 1 second, including time spent awaiting. Text extraction takes milliseconds for invoices and receipts, but long or image-heavy documents take longer. Use pdf-lib when you only need metadata.
- Memory. Each execution has 128 MB of memory, and parsing a large PDF uses much more memory than the size of the file.
- Body size. PDFs larger than 20 MiB arrive with
request.bodyset tonulland are delivered unchanged. Check fornullif you might receive large files. - Scanned PDFs. Text extraction only returns text stored in the PDF. Scanned documents are images and return no text, since transformations can't run OCR.
- Invalid PDFs. If the body isn't a valid PDF, unpdf throws
InvalidPDFException: Invalid PDF structure.and the transformation fails, which opens a transformation issue. To deliver those files unchanged instead, wrap the parsing intry/catchand return the request without changing its body.
For other modules and libraries you can use, see Transformation Node.js compatibility.