TAR archives are common in backups, software distribution, data pipelines, container workflows, and deployment systems. Python’s tarfile module makes these archives easy to read and extract, but extracting content from an external source is not a harmless file-copy operation. A crafted archive may attempt to write outside the destination directory, create dangerous links, preserve unwanted permissions, or overwrite sensitive files. The extraction_filter feature gives applications a clear way to define what is acceptable.
This guide explains how extraction filters work, how to choose a built-in policy, how to create a custom filter, and how to combine filtering with size limits, temporary directories, logging, and tests.
Why TAR extraction needs a policy
Each TAR member contains metadata such as its name, path, type, mode, owner, group, size, and link target. Those values come from the archive and should not automatically be trusted. A member named ../../settings.py attempts to leave the intended directory. An absolute path can point directly to a system location. A symbolic link may redirect later writes to a different location.
For that reason, treat an unknown TAR archive as structured input that must be validated. A safe workflow needs a dedicated destination, path checks, limits, an allowlist of accepted types, and a clear response when a member is rejected.
Using the data filter
The simplest modern approach is to pass a filter to the extraction operation. For archives containing ordinary application data, the data policy is usually the right starting point.
from pathlib import Path
import tarfile
archive = Path("backup.tar.gz")
destination = Path("extracted_data")
destination.mkdir(parents=True, exist_ok=True)
with tarfile.open(archive, "r:gz") as tar:
tar.extractall(destination, filter="data")
The filter is applied while members are being considered for extraction. It can reject unsafe entries or normalize metadata that should not be preserved in a data-oriented workflow.
Built-in policy choices
The data policy favors ordinary data extraction. The tar policy preserves more traditional archive behavior. The fully_trusted policy is permissive and should only be used when the archive is genuinely trusted, produced in a controlled environment, and required to retain broader TAR semantics.
For uploads, downloaded assets, partner integrations, or user-provided backups, prefer data. A more permissive option should be an explicit architectural decision, not a compatibility shortcut.
Creating a custom extraction filter
A project may need tighter rules. For example, a data import service may accept only CSV, JSON, and text files, reject links and devices, cap individual file sizes, and normalize permissions.
from pathlib import Path
import tarfile
ALLOWED = {".csv", ".json", ".txt"}
MAX_FILE = 20 * 1024 * 1024
def import_filter(member, path):
name = Path(member.name)
if member.isdir():
return member
if not member.isfile():
return None
if name.suffix.lower() not in ALLOWED:
return None
if member.size > MAX_FILE:
return None
return member.replace(mode=0o600)
with tarfile.open("incoming.tar") as tar:
tar.extractall("staging", filter=import_filter)
Returning None skips a member. Returning a member accepts it, optionally with adjusted metadata. This keeps the security policy close to the extraction code and makes it easier to review.
Reject links when they are unnecessary
Symbolic links and hard links are useful in system archives, but most data import jobs do not need them. Rejecting all links removes an important class of path-redirection problems.
def no_links(member, path):
if member.issym() or member.islnk():
return None
return tarfile.data_filter(member, path)
This example layers a project-specific rule over the standard data policy. Reusing the standard filter avoids reimplementing every validation detail yourself.
Limit the number and total size of members
An archive can exhaust resources even when each individual file is small. It may contain hundreds of thousands of members or declare a very large total uncompressed size. Check both dimensions before extraction.
MAX_MEMBERS = 5000
MAX_TOTAL = 500 * 1024 * 1024
with tarfile.open("incoming.tar") as tar:
members = tar.getmembers()
if len(members) > MAX_MEMBERS:
raise ValueError("Too many TAR members")
total = sum(item.size for item in members if item.isfile())
if total > MAX_TOTAL:
raise ValueError("Archive is too large")
tar.extractall("staging", members=members, filter="data")
Choose limits based on the application, available disk space, expected workloads, and operational timeouts.
Use a staging directory
Do not extract an untrusted archive directly into a live application directory. Create a fresh temporary directory, extract there, validate the resulting files, and move only approved content into the final location.
A staging workflow also makes cleanup easier. If validation fails, delete the entire temporary tree. The active data remains untouched and users never see a partially extracted result.
Validate content after extraction
Archive-level checks are only the first layer. A file with an allowed .json extension may still contain invalid or unexpected data. Parse structured files, validate schemas, enforce row counts, inspect encodings, and reject content that does not match the business contract.
For executable formats, templates, or configuration files, apply additional review. Avoid treating filename validation as content validation.
Error handling and cleanup
Do not silently ignore extraction errors. Record the operation identifier, archive source, rejected member, and policy reason. However, avoid exposing internal filesystem paths or stack traces to end users.
import tarfile
try:
with tarfile.open("incoming.tar") as tar:
tar.extractall("staging", filter="data")
except tarfile.TarError as exc:
raise RuntimeError("Invalid or unsafe TAR archive") from exc
Always clean the staging directory after failure. In long-running services, abandoned extraction directories become both a storage problem and an operational risk.
Run with minimal privileges
The extraction process should use an account that cannot write to system directories, application code, secrets, or unrelated user data. Filesystem permissions are an important second layer if an application-level check fails.
For highly untrusted workloads, isolate extraction in a container, sandbox, worker process, or disposable environment with CPU, memory, disk, and time limits.
Compatibility across Python versions
If your library supports multiple Python releases, verify the availability and behavior of the filter argument. Do not silently fall back to an unsafe path on an older interpreter. If extraction filtering is a core security requirement, declare a minimum Python version and enforce it during startup or installation.
Tests worth adding
Build fixtures that contain normal relative paths, parent-directory segments, absolute paths, symbolic links, hard links, oversized files, too many members, blocked extensions, unusual permissions, and malformed headers. Confirm that unsafe members are rejected and that the final destination remains unchanged after failure.
Also test valid archives from every producer your system supports. Security controls should be strict, but they should also be predictable and documented.
Operational logging
Useful metrics include accepted members, rejected members, declared total size, extraction duration, cleanup failures, and policy names. These signals help identify abuse, broken integrations, or an archive producer that changed unexpectedly.
Related Academify guides
Continue with Python zipfile.Path, Python os.path.splitroot, Python pathlib.Path.info, and Python sqlite3 autocommit. For authoritative details, read the Python tarfile documentation and the Python pathlib documentation.
Conclusion
tarfile extraction_filter turns extraction into a policy-driven operation. Start with filter="data" for external archives, reject unnecessary links and special files, enforce member and size limits, use a temporary destination, validate extracted content, and run with minimal privileges. These layers create a safer and more maintainable TAR processing pipeline in Python.







