Projects in Claude Code lets Claude split your work into threads inside a single conversation and run them in parallel in the cloud. I covered how it works and who can use it in What Claude Code Projects is. This article is the follow-up: a record of actually having it build an entire website.

I tried it from September 26 to 28, 2026, while Projects was still in public beta (rolling out gradually to Pro and Max; see the official documentation). The screens and behaviour may well change. What follows is what actually happened at the time, with numbers I checked against the screens, the usage report and the repository history.

The short version: what I learned by using it

Measured September 26–28, 2026 (the project's usage report and the repository history)

What it built

11 PRs

10 threads, about 23,000 lines. Work kept going while I slept.

Time taken

About 20 hours

From creation to the last merge, about 7 hours of it overnight.

Tokens used

About 190 million

97.7% were cache reads. Per unit of work, about the same as local Claude Code.

Result

An 80% first draft

Before launch I found gaps between the threads' responsibilities and bugs that only real data exposed.

In one line: development moves forward remarkably well with little supervision, but what comes out is a first draft, not a finished product, and deploying to a server you can only reach over SSH hits a wall. Below I go through the good and the bad in the order it happened.

1. What I had it build: the setup, and a confession up front

The subject was a database site for looking up how much memory local LLMs need. To answer "will this model run on my machine?", it collects quantized file sizes from Hugging Face, calculates the memory required at each context length, and lets you search backwards from the memory in your GPU or Mac. It is not too small, and it splits naturally into parts that can run in parallel (data collection, calculation, pages, reverse lookup, SEO), which made it a good test bed for Projects.

ItemThis run
What to buildA database of local LLM memory requirements (a Japanese-language site).
StackLaravel 13, PHP 8.5, MySQL 5.7. Hosted on shared hosting reachable only over SSH.
RepositoryA private repository on github.com. Project threads can only work with github.com repositories that have the Claude GitHub App installed, so I created a new GitHub account just for Claude.
PlanMax (20x).
ModelsThreads ran on Sonnet at medium effort by default; the coordinator picked Opus only for work where mistakes would hurt, such as reviews and calculations. The coordinator itself stayed on its default, Opus at low effort.

⚠️ A confession up front: in the first half I tied its hands

My first project instructions included rules that put human approval at every step: "propose each thread and wait for my OK before starting it", "no more than three threads at once", "a human approves every merge to main". I meant it as a safety measure, but it switched off the very thing that makes this feature worthwhile, letting Claude hand out the work and drive it forward. Partway through I rewrote the instructions to hand things over, and Section 4 compares the before and after. Some of the friction in the first half came from my instructions, not from the feature.

2. Getting started, and the five places I tripped

The steps themselves are short: in the desktop app's Code tab, choose Projects → New, enter a name, a goal and a repository, and create it. Around those steps, though, I tripped in these five places.

① GitHub App scope

The permission screen has "All repositories" selected by default. Unless it is a dedicated account, narrow it to "Only select repositories", because threads can add other repositories owned by the same account on their own.

② It runs once the moment you create it

With your first project, Claude starts a thread that reads the repository as soon as the project is created. It runs before you can paste any instructions, which is why its suggested threads came out in English for me.

③ The default is Opus

Threads default to Opus (medium effort on my screen; the official documentation says high). It drains your limits fastest, so right after creating a project, go to Settings → General and review the thread model.

④ Network allowlist

Hugging Face and hardware makers' sites are not on the default allowlist. *.nvidia.com only matches subdomains, so nvidia.com itself needed a separate line.

⑤ Changes don't reach running threads

Changes to the environment or the project instructions only take effect from new threads (the official documentation says so too). For a thread that had stalled, I had the coordinator continue the work in a new thread.

On ②: right after creation, the conversation showed the notice "Up to $100 of initial usage, including the automatic setup, won't count towards your usage limits". In other words, the first $100 worth of usage is not counted against your normal limits, and the usage screen also listed it as a "project setup credit" (about 24 hours until it expired). As of September 28, this perk is not mentioned on the Projects page of the official documentation. Section 6 shows how quickly it actually ran down.

Where I got lost in the environment settings was editing an existing environment. Going in through "Add cloud environment" opens a screen for a new environment, and at first I typed the allowed domains into the setup script field. To edit an existing environment, hover over it in the list and click the gear icon that appears (that is exactly what the official documentation says, but it is hard to discover from the screen alone).

3. About 20 hours, logged: what happened while I slept

Here is the timeline from creation (around 22:00 on September 26) to the last PR being merged (around 17:30 the next day, the 27th). All times are Japan Standard Time (JST).

26th, 22:10–23:50  Foundation and review

The foundation thread (Sonnet) created the Laravel skeleton, the database design and a rules document, and opened a PR. When I had a separate thread review it on Opus, it spun up MySQL 5.7 in a container, tried it for real, and found a bug where a timestamp column meant for record-keeping was overwritten with the current time every time a row was updated (MySQL 5.7 attaches auto-update to the first TIMESTAMP column, which never shows up in tests with sample data). A CI proposal was drafted at the same time, and I approved and merged PR #1.

27th, 0:00–7:30  Three threads in parallel overnight

Data collection (Opus), memory calculation (Opus), and SEO plus the shared layout (Sonnet) ran at the same time, and by morning all three were "awaiting review". What struck me was that the threads coordinated with each other through the coordinator: the calculation thread asked how to use the layout, and the layout thread answered. A problem the calculation thread found, that config files can't be fetched for models that require accepting terms of use (401), was passed on to the data collection thread, which dealt with it.

27th, 7:30–8:50  Cleaning up conflicts

Because the threads had been editing the same file (the route definitions) in parallel, the next PR conflicted after the first one was merged. The first time, the coordinator noticed the merge and told the thread to resolve the conflict on its own initiative. The second time the coordinator did nothing; a "Resolve conflicts" button appeared on the thread's card, and pressing it started the resolution. It does not react the same way every time.

27th, 8:50–11:30  Stopping at an unreachable site

The thread adding GPU and Mac hardware data couldn't reach the manufacturers' official sites from the cloud, and instead of filling in guesses, it stopped and showed a card with three options (widen the allowlist / have a human supply the values / look it up on the human's PC). After I fixed the allowlist and had it continue in a new thread, it added 25 models from the official pages. Along the way, the thread itself noticed that the page-summarising tool had invented a product name that doesn't exist, and from then on it switched to checking the pages' raw HTML directly.

27th, 17:00–17:30  Handed over, it ran on its own

When I rewrote the project instructions to hand things over, the coordinator announced, without my saying a word, "I've read the new instructions on how to work. From here on I'll decide the next tasks and move them forward", and went on to decide everything from merging the remaining PRs to adding more hardware data (two threads) by itself.

One thing that did not go well also deserves a mention. While building the data collection foundation, a thread hit the Hugging Face API 520 times in a row to work out how the external service's rate limiting behaves. The review thread judged this "not reasonable" and wrote up rules for using external APIs, and the later data collection thread made only 8 requests over its entire run. Left alone, it won't think about the load it puts on outside services, so it is worth spelling that out in the instructions.

4. Approving every step versus handing it over

As I said in Section 1, the first half ran on instructions that put human approval at every step, and at 17:00 on the 27th I rewrote them to hand things over. The behaviour changed clearly.

SituationFirst half: approval at every stepSecond half: handed over
Starting threadsThe first time, it ignored "propose and wait" and started right away. Once I reinforced it with a message saying "don't start until I say OK", it complied.The coordinator decided and started them itself.
MergingA human pressed the button on GitHub every time (7 times).Threads merged PRs themselves once CI passed (4 times).
Next taskA human decided and asked for it.The coordinator picked from the TODO list and started a new thread.
Human involvementMerges, conflict buttons, environment settings and relaying messages meant a lot of switching between screens.Almost none (only when environment settings were needed).

There were two things to watch out for when handing over. The first is that rewritten instructions don't reach threads that are already running. The coordinator itself explained, "This thread started before the instructions were rewritten, so it can't see the new instructions." The second is that threads couldn't delete old rules left in the project's memory. A task that tried to rewrite an old note saying "only a human merges" was blocked by a safety check. A human has to delete old notes from Settings → Memory.

Here is the skeleton of the instructions I eventually settled on.

This project builds and runs [your site].
How to proceed is up to you: the coordinator may decide which threads to start, how many, which models, and in what order.
You may merge PRs once CI passes. Resolve conflicts yourself.

Only ask a human about:
- Anything that costs money (paid APIs, paid services)
- Anything that needs secrets (API keys, passwords)
- Deploying to production or changing production data
- Decisions where the options diverge and either would be reasonable
- Anything you need but can't reach (don't fill gaps with guesses or dummy data)

Respect rate limits on external APIs, and don't hit them heavily for investigation.

That said, given what I learned afterwards, I recommend adding "changes to CI configuration and deployment scripts need human approval". Handed-over threads can change CI configuration too, and combined with a deployment mechanism that creates a path for changes to reach production without anyone checking (Section 7).

5. The quality I found after launch: every test had passed

I took what the project had built, handed it over to my usual local Claude Code, and put it into production. My first impression right after launch was "this is kind of... ordinary". There wasn't a single model page. Tracking down why revealed a gap unique to parallel development.

A gap between responsibilities: nobody decided "can this be published?"

Data collection thread

Wrote in its PR that "deciding whether something can be published is the page thread's job" and never built that check

Page thread

Built only the "don't show anything not publishable" side

Result

Nothing anywhere set the "publishable" flag, so no matter how much data went in, zero models were shown

Ten threads, well over a hundred tests, Opus reviews and CI: none of them caught this omission, because every thread was correct within its own responsibility. Without a role that looks at the whole thing end to end, gaps open up at the boundaries between tasks.

Loading real data surfaced two more problems.

  • The quantization tables included files that weren't the main model. Auxiliary models for speculative decoding (MTP, EAGLE and so on) and LoRA files were being counted as quantizations of the main model, so the table for one 12B model opened with a row reading "Q8_0 0.47GB" (the real Q8_0 is 12.7GB). After the fix, 195 files across 56 repositories were reclassified as auxiliary.
  • The newest, most in-demand models showed memory requirements as "cannot calculate". The formula as originally written didn't support newer-generation architectures (such as ones where different layers keep memory in different ways). The site's main selling point was missing on exactly the pages people view most.

Both came to light only once real data went into production, because the tests were built entirely on sample data. Here is what I fixed with local Claude Code, and how long it took (17:21 to 22:54 on September 28, 13 commits).

What was fixedHow it was found
There was no privacy policy page (required before running ads)Review by the server-administration role
The publishing check (linking models to their families) was missing entirelyZero models in production
sitemap.xml returned an error (500) in production; error notification emails couldn't be sentChecking in production
The contact form threw an error (500) on input in a different character encodingChecking in production
Auxiliary model files mixed in; memory requirements for newer architecturesEyeballing the pages with real data

There were clear positives too. Pages were fast (around 0.1 seconds for the main ones), file sizes were Hugging Face's actual values, and memory requirements for supported models matched the formula. SEO work such as titles, structured data, the sitemap and llms.txt was in there from the start. My overall verdict: the skeleton and the details were well made; what was missing was the "joints" between parts and real data.

6. Usage and cost: where 191.8 million tokens went

The Usage screen in the project settings shows tokens by thread and by model. A button in the top right copies it all as text. The report as of 11:47 on September 27 looked like this.

ItemValue
Threads10
Total tokens191.8 million (input 317K / output 574K / cache reads 187.4 million / cache writes 3.5 million)
Cache hit rate98%
Code changes+23,531 lines / −302 lines (the 8 threads that opened PRs)
Coordinator3.3 million (2% of the total)
Heaviest threadSEO and shared layout (Sonnet): 50.5 million (26%)

190 million sounds like a lot, but 97.7% of it was cache reads. A thread rereads the conversation so far on every action, so when one thread takes hundreds of actions, this is the shape you get. The coordinator added only 2%, so the management overhead was small.

For comparison, I took "tokens read ÷ tokens output" and set it against local Claude Code (the last three days on my own PC): about 330 for the project and about 340 locally. Consumption per unit of work was roughly the same as a local session. Projects feels heavier most likely because everything runs in parallel at once, so usage drains in a short, concentrated burst (the work itself differed, so treat this as a rough comparison).

On the cost side, I was able to track how the setup credit ($100) ran down.

Project setup credit ($100) used

Right after creation (the automatic run)1%
After the foundation6%
After the three overnight threads32%
After five parallel tasks78%
Extra work after handing overUsed up (100%)

Source: usage display in the desktop app (September 26–27, 2026). My weekly Max usage did not go up during this period.

This credit behaved much like a dollar amount counted at API prices. Pricing the 11:47 report at Anthropic's official rates (Pricing: Sonnet 5 at $2 input, $10 output and $0.20 cache reads; Opus 5.5 at $4 input, $20 output and $0.20 cache reads, all per million tokens) comes to about $58–65, which roughly matches what the screen showed at the time (78% = $78). So "$100 worth" means the amount that would cost $100 if you paid through the API, which is not that much in the context of a flat-rate Max plan. A separate credit handed out for cloud sessions ($250 on Max) said "Projects are not eligible" on its claim screen, and indeed not a single dollar of it was used.

7. Why deploying to production got stuck

The hardest part was getting what had been built into production, because the host was shared hosting reachable only over SSH.

  • Cloud threads can't reach the production server (the SSH key lives only on my local PC).
  • To run on the local PC you use the project's "Work locally" option. Under the hood it is Remote Control, and it requires turning on "Use this computer from your phone and claude.ai" in the desktop app.
  • However, that setting applies to the list of every folder you have opened in Claude Code so far, gathered automatically. In my environment there were 22, one of which was a parent folder containing dozens of projects. While it is on, all of them become places where work can be started remotely, and the folder names, paths and repository URLs are sent to Anthropic as well. To limit it to a single folder, you have to use a different route: open a terminal in that folder and run claude remote-control.

So I weighed several ways to deploy.

MethodHow it worksAssessment
Work locally (Remote Control)A thread running on the local PC deploys over SSHWidens the set of exposed folders. Toggling it on and off every time is a chore.
A runner resident on the PC (self-hosted GitHub Actions)When main is updated, the deployment runs on the local PCIf threads are allowed to edit and merge workflows, it becomes an entry point for running arbitrary code on the local PC. Rejected.
Notify the server via webhookThe server receives a notification from GitHub and pullsMeans opening a new URL that anyone outside can hit. Dropped after review by the server-administration role.
The server pulls on a scheduleA cron job on the server checks GitHub, fetches with a read-only key and applies the updateNo inbound entry point, so it is safe. But by this point I had decided to move to my usual local workflow.

In the end, I archived the project and moved what it had built into my usual local Claude Code workflow. I judged that putting it on the same deployment pipeline as my other sites was safer than adding a new mechanism.

Looking back, it seems better to think of Projects as designed to pair with hosting that deploys automatically when you merge on GitHub. With Vercel, for example, connecting it to GitHub is enough for changes to flow all the way to production, so the kind of dead end I hit would not happen. That said, Vercel's free Hobby plan is limited to non-commercial use, and running ads such as Google AdSense requires the paid Pro plan (from $20 a month) (Fair Use Guidelines). It is also not a natural fit for a PHP and MySQL site like this one. The key is to decide the stack and the host together, up front.

I plan to cover the risks of letting your own PC be operated remotely, and how that differs from everyday Claude Code, in a separate article. Using your local session from a phone is covered in the Remote Control article.

8. What it is good for, and what it is not

Good fit

  • Something new that splits into independent parts
  • Work you want to move forward while you sleep or are out
  • A host that deploys automatically through a GitHub integration
  • You can write down up front how much to hand over

Poor fit

  • Servers that need SSH to deploy, using a key on your own machine
  • Existing work that depends heavily on local checking tools or custom procedures
  • Repositories hosted somewhere other than GitHub
  • Small jobs that finish in one session (a cloud session is enough)

Based on this experience, here is what I would check before starting.

  • Host: does merging on GitHub deploy automatically? If SSH is required, decide in advance on something like having the server pull.
  • GitHub permissions: the scope of the Claude GitHub App install. Use a dedicated account, or narrow the repositories.
  • How much to hand over: in the project instructions, list only what should be asked of a human. Make changes to CI and deployment configuration require approval.
  • Network: if you use external APIs or sites, put both the bare domain and its subdomains on the allowlist.
  • Checking with real data: don't relax just because the sample-data tests pass. At the end, start one thread whose job is to push production-like data through end to end and look at the result.
  • Usage limits: review the default thread model. For the first 24 hours there is a $100 setup credit.

Summary

Projects genuinely takes over the work of a human handing out tasks, chasing them, and re-explaining the same background. Three tasks ran in parallel overnight, the threads coordinated with each other, and once handed over it made its own calls on merges and what to do next. Tokens per unit of work were no different from local Claude Code.

On the other hand, what comes out is a first draft. Splitting work in parallel opens gaps at the boundaries between responsibilities. Every sample-data test can pass and real data will still expose bugs. And it does not get along with servers you can only reach over SSH. If you use it, three things should spare you the detours I took: choose a host with a GitHub integration, write down how much to hand over at the start, and at the end start one thread whose job is to check everything end to end with real data.

FAQ

Q. How much does Projects cost?

A. There is no separate charge; it draws on your normal Pro or Max usage limits. On top of that, this time I got a "setup credit" under which up to $100 of usage in roughly the first 24 hours didn't count toward my limits (shown on screen; not in the official documentation as of September 28). The work, 10 threads and 11 PRs, used about 190 million tokens and used up that credit. At API prices, that is roughly $100 worth.

Q. Does it really keep going when I close my laptop?

A. Yes. Threads running in the cloud kept going whether I closed the PC or put it to sleep. In this run, three tasks finished during about 7 hours overnight. Threads started on your own PC with "Work locally", however, only run while the PC is awake.

Q. Can I use it with repositories outside GitHub?

A. Threads that work with code assume a github.com repository with the Claude GitHub App installed. I was managing this project on my own git server, so I created a GitHub account just for Claude, started there, and moved everything back to my local setup at the end.

Q. What happens if I change the project instructions partway through?

A. The coordinator picks them up immediately, and its way of working changed without my having to say anything. They don't reach threads that are already running, though, so if needed, have the work continue in a new thread. Old rules left in the project's memory had to be deleted from Settings by a human.

Q. Is "Work locally" safe?

A. The only traffic is an encrypted outbound connection from your PC; no listening port is opened. However, turning on the setting in the desktop app makes the list of folders gathered from your usage history (22 in my case, including dozens of projects inside a parent folder) into places where work can be started remotely. You need practices such as turning it on only while you use it and making sure the account you sign in with has two-factor authentication.

Q. Can I use what it builds as is?

A. In my case, no. The skeleton and the SEO work were well done, but the logic at the boundaries between responsibilities was missing entirely, and there were bugs that only showed up with real data. Before launch, you need a step that loads real data and checks the whole thing end to end.