File indexer
The file indexer indexes files from the file collections selected in the indexing service. For PDF files it also extracts the text content.
Excluding a file
The extension adds an Include in Search option to the file metadata, on a new Behaviour tab. It is enabled by default. Disable it to keep a file out of the index.
Exclude a file from indexing
Fields
In addition to the standard fields, the file indexer sends:
| Field | Description |
|---|---|
extension | The file extension. |
mime | The MIME type of the file. |
name | The file name. |
size | The file size in bytes. |
url | The public URL of the file, without the leading slash for files in a local storage. |
content | The extracted text, only for PDF files. |
The mapped metadata fields and the allowed file extensions default to:
module.tx_typo3searchalgolia.indexer.sys_file_metadata {
fields {
title = title
description = description
alternative = alternative
creator = author
}
# Comma-separated list of allowed file extensions
extensions = pdf
}
Indexed file content becomes searchable. The protection of non-public storages is not evaluated. Their files are indexed like public ones. Keep sensitive documents out of the indexed file collections, or exclude them individually.