Path.walk() is a modern way to traverse directory trees with Python’s pathlib module. It yields the current directory, child directory names, and file names while keeping the current location as a Path object. This makes directory automation easier to read, compose, and maintain across operating systems.
How Path.walk works
Each iteration returns three values: the current directory, a list of subdirectory names, and a list of file names. Join a returned name with the current directory by using the / operator.
from pathlib import Path
root = Path("project")
for directory, dirnames, filenames in root.walk():
for name in filenames:
path = directory / name
print(path)Why pathlib is useful
Path objects provide properties and methods such as suffix, stem, name, stat(), exists(), and is_file(). This avoids manual string concatenation and platform-specific path separators.
Related Academify guides include importlib.resources, fileinput, linecache, and contextlib.chdir.
Processing files as a stream
For large trees, process files as they are discovered instead of storing every path in memory. This streaming approach scales better and lets the application apply filters early.
for directory, dirnames, filenames in Path("data").walk():
for name in filenames:
process(directory / name)Filtering by extension
Use suffix.lower() to make extension checks predictable. A set is convenient when several extensions are accepted.
allowed = {".csv", ".json", ".parquet"}
for directory, dirnames, filenames in Path("data").walk():
for name in filenames:
path = directory / name
if path.suffix.lower() in allowed:
process(path)Top-down and bottom-up traversal
With top_down=True, parents are visited before children. This is the default and allows pruning before Python enters a child directory. With top_down=False, deeper directories are visited first, which is useful for reports that need child information before parent information.
Skipping directories
When walking top down, modify dirnames in place to prevent traversal into locations that are not relevant.
ignored = {".git", ".venv", "node_modules", "__pycache__"}
for directory, dirnames, filenames in Path("project").walk():
dirnames[:] = [name for name in dirnames if name not in ignored]
for name in filenames:
print(directory / name)Assigning to the slice matters because the traversal engine uses the original list. Rebinding a local variable does not prune the walk.
Error handling
Permission failures, directories changed by another process, unavailable mounts, and input/output errors can occur during traversal. The on_error callback receives an exception and lets the application log the problem or stop with a clear message.
def report(error):
print(f"Cannot access {error.filename}: {error}")
for directory, dirnames, filenames in Path("workspace").walk(on_error=report):
passDo not silently discard every error. An incomplete inventory can be misleading in backup, migration, and compliance workflows.
Symbolic links and cycles
Following directory symbolic links can move the traversal outside the intended root or create cycles. A link may point to an ancestor and repeat the same area. Before following links, define a visited-path strategy and decide whether external locations are allowed.
Calculating total size
You can combine walk() with stat() to estimate the total size of a directory tree.
def total_size(root: Path) -> int:
total = 0
for directory, dirnames, filenames in root.walk():
for name in filenames:
path = directory / name
try:
info = path.stat()
total += info.st_size
except OSError:
continue
return totalThe result is not atomic. Files may change, disappear, or grow while the scan runs. Treat it as an operational estimate unless the storage system provides a snapshot.
Finding large files
A common maintenance task is identifying files above a threshold. A safer workflow is to produce a report first, review the paths, and only then decide what action is appropriate.
limit = 500 * 1024 * 1024
for directory, dirnames, filenames in Path("archive").walk():
for name in filenames:
path = directory / name
try:
if path.stat().st_size > limit:
print(path)
except OSError as error:
print(error)Path.walk versus rglob
rglob() is concise for pattern searches such as *.py. walk() is better when you need pruning, explicit error handling, traversal order, or decisions based on the current directory.
for path in Path("project").rglob("*.py"):
print(path)Path.walk versus os.walk
os.walk() remains reliable and widely compatible. Path.walk() mainly improves integration with the object-oriented pathlib API. Existing code does not need an immediate rewrite, but new code may become clearer when paths remain Path objects from discovery through processing.
Limiting depth
The method has no direct maximum-depth argument, but you can calculate depth relative to the root and clear dirnames when the limit is reached.
root = Path("data")
maximum = 3
for directory, dirnames, filenames in root.walk():
depth = len(directory.relative_to(root).parts)
if depth >= maximum:
dirnames.clear()Performance practices
Call stat() only once per file when possible. Prune ignored directories early. Process items in a streaming fashion. Avoid expensive checks when a simple suffix or name comparison works. On network file systems, metadata calls can be much slower than local operations.
Safe file operations
Directory traversal often supports maintenance tools. Before moving, replacing, or cleaning files, validate that each path remains inside the approved root. Separate discovery from execution, produce logs, and make the operation reviewable. Symbolic links and concurrent filesystem changes require extra care.
Compatibility
Path.walk() is available only in newer Python versions. Check the minimum interpreter supported by your project. For older environments, use os.walk() or a compatibility wrapper.
Review the official pathlib documentation and the official os.walk documentation for current details.
Designing a reusable scanner
A robust file-processing tool usually separates traversal, selection, processing, and reporting. The traversal layer discovers paths. The selection layer applies extension, size, or naming rules. The processing layer performs the intended work. The reporting layer records successes, skipped paths, and errors.
This separation makes testing easier. You can test path selection without changing files, replace the processing function with a mock, and confirm that errors remain visible.
Testing traversal code
Use temporary directories in automated tests. Create nested folders, ignored names, mixed-case extensions, and changing files. Where supported, test symbolic links and permission failures. Verify that the traversal returns exactly the expected set and does not enter excluded locations.
Common mistakes
Common mistakes include treating returned file names as complete paths, forgetting to modify dirnames in place, calling stat() repeatedly, ignoring all exceptions, following links without cycle protection, and assuming the directory tree remains unchanged during the entire walk.
Conclusion
Path.walk() provides an expressive interface for traversing directory trees. It works especially well with early pruning, explicit error handling, careful symbolic-link policies, and streaming processing. With these practices, it becomes a strong foundation for inventories, backups, data pipelines, validators, and maintenance tools.







