Project

Git from Rust

A from-scratch Git implementation in Rust, covering content-addressed storage, trees, commits, packfiles, deltas, and cloning.

August 2026 — September 2026

RustGitSystemsNetworkingVersion Control

What actually happens when Git stores a file, creates a commit, or clones a repository?

I built a small Git implementation from scratch in Rust as part of CodeCrafters' Build Your Own Git challenge.

The project started with git init and gradually turned into an implementation of Git's content-addressed object model, including blobs, trees, commits, object compression, packfiles, delta resolution, and repository cloning.

View the source on GitHub →

What I built

🧱 Git's Object Model

Implemented Git's core object types:

  • Blobs
  • Trees
  • Commits

Objects are addressed by their SHA-1 hash and stored under .git/objects/, following Git's content-addressed storage model.

The same mechanism is used to hash, compress, store, and retrieve objects.

🔐 Hashing & Object Storage

Implemented:

  • hash-object
  • SHA-1 object hashing
  • Git object headers
  • Zlib compression
  • Reading compressed objects back from .git/objects

A Git object isn't simply the file contents.

Its hash is calculated from:

<type> <size>\0<content>

That small detail makes Git's storage model much more interesting than simply hashing files.

🌳 Trees & Snapshots

Implemented write-tree to recursively construct a tree representing the current working directory.

The implementation:

  • Walks directories recursively
  • Creates blob objects for files
  • Creates tree objects for directories
  • Encodes object hashes as raw 20-byte SHA-1 values
  • Sorts tree entries
  • Stores the resulting tree object

This makes the connection between a filesystem and Git's snapshot model explicit.

📜 Commits

Implemented commit-tree with:

  • Tree references
  • Parent commits
  • Commit messages
  • Author information
  • Committer information
  • SHA-1 object creation

A commit is therefore just another immutable Git object pointing at a tree and, optionally, its parent commit.

That naturally forms Git's history graph.

🔍 Reading & Inspecting Objects

Implemented:

  • cat-file
  • ls-tree

cat-file decompresses an object from .git/objects and retrieves its contents.

ls-tree parses the binary tree representation and reconstructs the entries stored inside it.

📦 Packfiles

The more interesting part came when implementing clone.

Git doesn't simply send every object as an individual loose object. Repositories can transfer objects using packfiles.

The implementation can:

  • Request a packfile from a Git server
  • Parse the packfile header
  • Read individual object headers
  • Decompress packed objects
  • Identify object types
  • Store the resulting objects locally

🔀 Delta Objects

Packfiles can store objects as deltas rather than complete copies.

I implemented delta resolution, including:

  • Variable-length size encoding
  • Copy instructions
  • Insert instructions
  • Base-object references
  • Reconstructing the complete object from a delta

This was one of the more interesting parts of the project because it turns an apparently simple "download the repository" operation into a binary parsing and reconstruction problem.

🌐 Git HTTP Protocol

The clone implementation also communicates directly with a Git server.

It uses:

  • info/refs
  • git-upload-pack
  • HTTP requests
  • Git's packet-line format
  • Packfile responses

The implementation discovers the repository's current commit, requests the corresponding objects, reconstructs the packfile contents, and stores them in the local repository.

📂 Clone & Checkout

Implemented a basic clone flow:

  1. Create a new repository directory
  2. Initialize .git
  3. Discover the remote commit
  4. Request the packfile
  5. Parse and store objects
  6. Resolve delta objects
  7. Read the commit
  8. Traverse its tree
  9. Recreate the working directory

The final step recursively walks tree objects and reconstructs files from their blob contents.

What I learned

The interesting part of Git isn't the command line interface. It's the data model underneath it.

A file becomes a blob.

A directory becomes a tree containing references to blobs and other trees.

A snapshot becomes a commit pointing at a tree.

Commits point to other commits.

And everything is identified by its content.

That means Git's history isn't a collection of copied folders. It's a graph of immutable, content-addressed objects.

Building the storage format, tree construction, commit objects, packfiles, and delta resolution from scratch made that model much easier to understand.

Tech

Rust · SHA-1 · Zlib · Binary parsing · File systems · HTTP · Git protocol · Packfiles

The implementation does not use Git itself to perform these operations.

Challenge

This project is based on CodeCrafters' Build Your Own Git, a hands-on systems challenge that progressively implements parts of Git from scratch.

Try the challenge yourself →

Read the implementation →