# Extract Text and Data from a PDF with a Transformation

When a [Source](/docs/sources) receives an `application/pdf` request, a [Transformation](/docs/transformations) receives the PDF as a `Buffer` in `request.body`. You can parse it with an npm library and replace the body with JSON, so the [Event](/docs/events) becomes searchable and filterable and your [Destination](/docs/destinations) receives structured data instead of a file. The parsing runs in Hookdeck, so you don't need to run a document parser in your own service.

This guide extracts an invoice number and total from the text of a PDF with [unpdf](https://github.com/unjs/unpdf), and reads document metadata such as the title, author and page count with [pdf-lib](https://pdf-lib.js.org/).

## Prerequisites

* A [Connection](/docs/connections) whose source receives PDF files with `Content-Type: application/pdf`
* Node.js and npm on your machine, to bundle the transformation

Transformations can only import Node.js built-in modules, so you bundle the library into your code with [esbuild](https://esbuild.github.io/). See [bundle dependencies](/docs/transformations#bundle-dependencies) for background.

## Extract text and invoice fields

### 1. Install the dependencies

```bash
mkdir pdf-transformation && cd pdf-transformation
npm init -y
npm install unpdf
npm install --save-dev esbuild

```

### 2. Write the transformation

Save this as `pdf-text.js`. It extracts the text of every page, then reads the invoice number and total with regular expressions. Adjust the expressions to the layout of your documents.

```js
import { extractText, getDocumentProxy } from 'unpdf';

export default {
  async transform(request) {
    const pdf = await getDocumentProxy(new Uint8Array(request.body));
    const { totalPages, text } = await extractText(pdf, { mergePages: true });

    const invoiceNumber = text.match(/Invoice number:\s*(\S+)/)?.[1] ?? null;
    const total = text.match(/Total due:\s*\$([\d,]+\.\d{2})/)?.[1];

    request.body = {
      invoice_number: invoiceNumber,
      total: total ? Number(total.replace(/,/g, '')) : null,
      pages: totalPages,
      text,
    };
    request.headers['content-type'] = 'application/json';
    return request;
  },
};

```

`new Uint8Array(request.body)` copies the bytes into a standalone `Uint8Array`, the type unpdf expects. A `Buffer` can be a view into a larger shared `ArrayBuffer`, and the copy keeps the parser from seeing anything outside the file.

### 3. Bundle it

```bash
npx esbuild pdf-text.js --bundle --format=esm --platform=node --minify --outfile=dist/pdf-text.js

```

The bundle includes the PDF parser and is about 1.6 MB minified, within the 5 MB code limit.

### 4. Add it to your connection

Paste the contents of `dist/pdf-text.js` into a new transformation on your connection. See [create a transformation](/docs/transformations#create-a-transformation). You can also upload the file with the [Hookdeck CLI](/docs/cli#hookdeck-gateway-transformation-upsert):

```bash
hookdeck gateway transformation upsert pdf-text --code-file dist/pdf-text.js

```

### 5. Send a PDF

```bash
curl -X POST "https://hkdk.events/src_xxxxxxxx" \
  -H "Content-Type: application/pdf" \
  --data-binary @invoice.pdf

```

Use `--data-binary`, not `-d`: `-d` strips newlines and corrupts binary files.

### What your destination receives

For a two-page invoice, your destination receives `Content-Type: application/json` and a body like this one:

```json
{
  "invoice_number": "INV-1042",
  "total": 1250,
  "pages": 2,
  "text": "ACME Corp\nInvoice number: INV-1042\nInvoice date: 2026-10-01\nBill to: Jane Doe\nWidget x 2 $500.00\nSupport plan $250.00\nTotal due: $1,250.00\nPayment terms: Net 30"
}

```

Because the event is JSON, you can [filter](/docs/filters) on `total` or search for `invoice_number` in the dashboard.

## Read document metadata

To read metadata without extracting text, use pdf-lib. It's smaller than unpdf and faster for large documents.

```bash
npm install pdf-lib

```

```js
import { PDFDocument } from 'pdf-lib';

export default {
  async transform(request) {
    const pdf = await PDFDocument.load(request.body, { updateMetadata: false });

    request.body = {
      title: pdf.getTitle() ?? null,
      author: pdf.getAuthor() ?? null,
      page_count: pdf.getPageCount(),
      created_at: pdf.getCreationDate()?.toISOString() ?? null,
      size: request.body.length,
    };
    request.headers['content-type'] = 'application/json';
    return request;
  },
};

```

```bash
npx esbuild pdf-metadata.js --bundle --format=esm --platform=node --minify --outfile=dist/pdf-metadata.js

```

For the same invoice, your destination receives:

```json
{
  "title": "Invoice INV-1042",
  "author": "Acme Billing",
  "page_count": 2,
  "created_at": "2026-10-08T21:21:22.000Z",
  "size": 1377
}

```

`updateMetadata: false` keeps pdf-lib from changing the producer and modification date as it loads the document. Fields the PDF doesn't set come back as `null`.

## Keep the original PDF

Replacing the body with JSON means your destination no longer receives the PDF. If it needs both, include the file in the JSON as base64:

```js
import { extractText, getDocumentProxy } from 'unpdf';

export default {
  async transform(request) {
    const original = request.body;
    const pdf = await getDocumentProxy(new Uint8Array(original));
    const { text } = await extractText(pdf, { mergePages: true });

    request.body = {
      invoice_number: text.match(/Invoice number:\s*(\S+)/)?.[1] ?? null,
      pdf_base64: original.toString('base64'),
    };
    request.headers['content-type'] = 'application/json';
    return request;
  },
};

```

Base64 makes the file about a third larger. Other options:

* Add a second connection on the same source without the transformation, so one destination receives the PDF and another receives the JSON.
* Leave `request.body` unchanged and put the extracted fields in headers, such as `request.headers['x-invoice-number']`. Returning the body unchanged keeps the original bytes.

## Limits and pitfalls

* Execution time. Transformations must finish within 1 second, including time spent awaiting. Text extraction takes milliseconds for invoices and receipts, but long or image-heavy documents take longer. Use pdf-lib when you only need metadata.
* Memory. Each execution has 128 MB of memory, and parsing a large PDF uses much more memory than the size of the file.
* Body size. PDFs larger than 20 MiB arrive with `request.body` set to `null` and are delivered unchanged. Check for `null` if you might receive large files.
* Scanned PDFs. Text extraction only returns text stored in the PDF. Scanned documents are images and return no text, since transformations can't run OCR.
* Invalid PDFs. If the body isn't a valid PDF, unpdf throws `InvalidPDFException: Invalid PDF structure.` and the transformation fails, which opens a [transformation issue](/docs/transformations#transformation-issues). To deliver those files unchanged instead, wrap the parsing in `try`/`catch` and return the request without changing its body.

For other modules and libraries you can use, see [Transformation Node.js compatibility](/docs/transformations#nodejs-compatibility).