When preparing reports or organizing project files, you may need to combine several PDFs into a single document. Sometimes you need all the pages from each file. Other times, you only want to include certain pages. If the PDF data comes from memory or a network response, working with input streams may be more convenient.
This article covers three ways to merge PDFs in Java: combining entire documents, merging selected pages from multiple PDFs, and merging PDFs using input streams.
1. Add the Maven Dependency
First, add the following repository and dependency to your Java project's pom.xml file:
<repositories>
<repository>
<id>com.e-iceblue</id>
<name>e-iceblue</name>
<url>https://repo.e-iceblue.com/nexus/content/groups/public/</url>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>e-iceblue</groupId>
<artifactId>spire.pdf</artifactId>
<version>12.9.0</version>
</dependency>
</dependencies>
For the examples below, we'll use three PDF files:
pdfs/
├── 01-cover.pdf
├── 02-report.pdf
└── 03-appendix.pdf
These files contain a cover page, the main report, and an appendix, respectively. Each example is a standalone Java program that saves the merged PDF to a separate output file.
2. Merge Multiple PDF Files into One
If you want to combine all pages from multiple PDF files, you can use the PdfDocument.mergeFiles() method.
This method accepts an array of PDF file paths and saves the merged document to a specified location.
Here's the complete Java code:
import com.spire.pdf.PdfDocument;
public class MergePDF {
public static void main(String[] args) {
// Specify the PDF files to merge
String[] files = {
"pdfs/01-cover.pdf",
"pdfs/02-report.pdf",
"pdfs/03-appendix.pdf"
};
// Specify the output file path
String outputFile = "merged.pdf";
// Merge the PDF files
PdfDocument.mergeFiles(files, outputFile);
System.out.println("PDFs merged successfully: " + outputFile);
}
}
The main method used here is PdfDocument.mergeFiles(files, outputFile), which takes two arguments:
files: An array containing the paths of the PDF files to merge.outputFile: The path where the merged PDF will be saved.
The files are merged in the order they appear in the array. In this example, the resulting merged.pdf contains all pages from the cover, followed by the main report and then the appendix.
To change the order, simply rearrange the file paths in the array. There's no need to modify the source PDFs.
3. Merge Selected Pages from Multiple PDF Files
Sometimes you don't need every page from each PDF. For example, you might want to combine the cover page of a report, a few pages from its main content, and the first page of an appendix.
To do this, load the source PDFs and insert the required pages into a new document using two methods:
insertPage(): Inserts a single page from a source PDF.insertPageRange(): Inserts a continuous range of pages from a source PDF.
Suppose we want to merge the following pages:
| Source PDF | Pages to include |
|---|---|
| 01-cover.pdf | Page 1 |
| 02-report.pdf | Pages 2–4 |
| 03-appendix.pdf | Page 1 |
The resulting PDF will contain five pages in the specified order.
Here's the complete code:
import com.spire.pdf.PdfDocument;
public class MergeSelectedPDFPages {
public static void main(String[] args) {
// Create PDF document objects for the source files
PdfDocument cover = new PdfDocument();
PdfDocument report = new PdfDocument();
PdfDocument appendix = new PdfDocument();
// Create a new PDF document for the merged pages
PdfDocument mergedPdf = new PdfDocument();
try {
// Load the source PDF files
cover.loadFromFile("pdfs/01-cover.pdf");
report.loadFromFile("pdfs/02-report.pdf");
appendix.loadFromFile("pdfs/03-appendix.pdf");
// Check whether the required pages exist
if (cover.getPages().getCount() < 1
|| report.getPages().getCount() < 4
|| appendix.getPages().getCount() < 1) {
throw new IllegalArgumentException(
"One or more source PDFs do not contain the required pages.");
}
// Insert page 1 from the cover
mergedPdf.insertPage(cover, 0);
// Insert pages 2 through 4 from the report
mergedPdf.insertPageRange(report, 1, 3);
// Insert page 1 from the appendix
mergedPdf.insertPage(appendix, 0);
// Save the merged PDF
mergedPdf.saveToFile("merged-selected.pdf");
System.out.println("Selected PDF pages merged successfully.");
} finally {
// Release PDF document resources
mergedPdf.close();
cover.close();
report.close();
appendix.close();
}
}
}
There are two important details to keep in mind when working with page indexes.
Page Indexes Start at Zero
In the code, page indexes are zero-based. This means page 1 has an index of 0, page 2 has an index of 1, and so on.
For example:
mergedPdf.insertPage(cover, 0);
This inserts the first page of the cover PDF into the merged document.
insertPageRange() Includes the End Index
The following statement:
mergedPdf.insertPageRange(report, 1, 3);
Inserts pages with indexes 1, 2, and 3 from report, corresponding to pages 2–4. Both the start and end indexes are included.
Pages are added in the order the insertion methods are called, so you can control the final page order by changing the sequence of these operations.
Page indexes must also fall within the source document's page count. The example checks the available pages before inserting them to avoid referencing pages that don't exist.
If you need several nonconsecutive pages, you can call insertPage() multiple times with different page indexes.
4. Merge PDF Files Using Input Streams
In addition to merging PDFs by file path, you can merge them using input streams.
For example, in a web application, PDF data might come from uploaded files, network responses, or byte arrays stored in memory. In these cases, you may not need to save each source PDF as a temporary file before merging.
The PdfDocument.mergeFiles() method has an overload that accepts an InputStream[] array, allowing multiple PDF streams to be merged directly.
The following example uses FileInputStream to demonstrate this approach:
import com.spire.pdf.FileFormat;
import com.spire.pdf.PdfDocument;
import com.spire.pdf.PdfDocumentBase;
import java.io.FileInputStream;
import java.io.IOException;
import java.io.InputStream;
public class MergePDFByStream {
public static void main(String[] args)
throws IOException {
// Open input streams for the PDF files
try (InputStream stream1 = new FileInputStream(
"pdfs/01-cover.pdf");
InputStream stream2 = new FileInputStream(
"pdfs/02-report.pdf");
InputStream stream3 = new FileInputStream(
"pdfs/03-appendix.pdf")) {
// Store the input streams in an array
InputStream[] streams = {
stream1,
stream2,
stream3
};
// Merge the PDF streams
PdfDocumentBase mergedPdf =
PdfDocument.mergeFiles(streams);
try {
// Save the merged PDF
mergedPdf.save(
"merged-stream.pdf",
FileFormat.PDF);
System.out.println("PDF streams merged successfully.");
} finally {
// Release the merged document resources
mergedPdf.close();
}
}
}
}
Unlike the first example, this approach uses input streams rather than file paths.
The FileInputStream objects provide access to the source PDF data. These streams are placed in an InputStream[] array and passed to:
PdfDocument.mergeFiles(streams);
The method returns a PdfDocumentBase object representing the merged document. We then use its save() method to write the result to merged-stream.pdf.
The example uses Java's try-with-resources statement to manage the input streams. They are automatically closed when the block exits, including when an exception occurs. The merged PDF document is also closed in the finally block.
One distinction is worth noting: merging PDFs through input streams doesn't necessarily mean the entire operation takes place in memory, nor does it prevent you from saving the result to disk.
In this example, the input streams read local files, and the merged PDF is saved to disk.
If the PDF data is already available as byte arrays, you can use ByteArrayInputStream instead of FileInputStream to avoid writing the source data to temporary files.
5. Things to Consider When Merging PDFs
1. Page Order Depends on the Input Order
When merging entire PDFs or input streams, the source documents are processed in the order they appear in the input array.
When merging selected pages, the final order depends on the sequence of insertPage() and insertPageRange() calls.
2. Pages May Have Different Sizes
If your source PDFs contain pages of different sizes, such as A4 and A3, the merged document may retain those differences.
Merging PDFs does not automatically resize or reformat their pages. If you need a consistent page size throughout the document, you'll have to handle that separately.
3. Password-Protected PDFs Require the Correct Password
If a source PDF is protected by an opening password, a standard merge operation may fail because the file cannot be accessed directly.
If you know the password, you can load the document using loadFromFile(filePath, password) and then merge the required pages into a new PDF.
Top comments (0)