← Journal

Research hypothesis · Not a shipped feature

Can Unjiggle learn from everyone without seeing anyone’s home screen?

Thomas Chung · October 1, 2026

I built Unjiggle around a simple constraint: your home screen is yours. The app reads its layout over USB on your Mac. I don’t want to collect everyone’s app lists just to make better suggestions. But I do want the suggestions to get better. Can both things be true?

My hypothesis is that, if people choose to participate, a shared model trained on their Macs could improve plan recommendations without sending raw layouts to me. Your Mac would learn from your feedback; a server would combine protected updates from enough participating Macs; the result would be a better starting point for each person. That’s federated learning. It’s an idea I want to test, not a description of how Unjiggle works today.

Why the Mac, not the iPhone?

Unjiggle runs on a Mac connected to your iPhone by USB. The iPhone doesn’t run Unjiggle code, so in this proposal the Mac would be the learning client. That makes the idea more practical than asking every phone to do training. It does not make the privacy or engineering problems go away.

There are also three different things people tend to lump together. Federated analytics asks an aggregate question, like what share of volunteered plans received a useful rating. Federated evaluation tests a proposed model against labels that stay on participating Macs. Federated learning actually trains a shared model from protected updates. I’d try them in that order. Being able to count something privately is not evidence that a useful global model exists.

First, is there even a signal?

Today the app can optionally remember whether you found a preset plan useful and show liked presets first. That feedback stays on your Mac. It’s off by default, can be deleted, and doesn’t train a shared model. It also rates an entire plan, not individual app moves. A tap on “Apply” isn’t proof of good taste; silence after a change isn’t approval.

Before building a federation, I’d ask volunteers for separate, informed permission to evaluate that signal. On held-out plans, do coarse feedback and a few privacy-safe features predict an explicit useful/not-useful rating better than simply showing everyone the same order? If not, the clever infrastructure has nothing useful to learn.

The present signal cannot teach a model where to place each individual app. That would require different product interactions, different labels, and a fresh consent review. Nor would I treat a share, a rerun, or the absence of an Undo as a positive label. A missing rating is missing data, not a quiet yes.

The actual comparison

Suppose there is a signal. I’d compare three approaches on held-out, voluntary ratings: fixed ordering, a model that learns only on your Mac, and a shared model that also adapts on your Mac. The test is whether the shared approach improves useful-plan ratings over local-only learning, not just over doing nothing. Home screens are personal. If local-only wins or the shared advantage is too small or uncertain, we should stop at local-only.

I’d publish the number of participants and labeled plans, how many people declined or didn’t rate, uncertainty around the result, and the comparison method. I wouldn’t call observational usage proof that a model caused an improvement. We don’t have those results yet.

What I would build, in order

First, make sure people actually get from the website to a useful scan. Page visits and download clicks are not installs, and installs are not successful scans. The app doesn’t send usage telemetry today. I won’t add a hidden install beacon to manufacture a clean-looking funnel.

Next, use the optional local feedback for a genuinely local benefit and see whether volunteers can supply enough explicit labels for research. Participation in that research would be a separate choice from both local feedback and hosted AI. If it’s appropriate to study a central prototype with volunteered data, that would need its own clear permission and privacy-policy change; the current toggle authorizes neither a research upload nor federation.

Only if there’s a learnable signal and enough willing Macs would I pilot aggregate statistics with a minimum group size and tested protections. Evaluation comes after that. Training comes last, and only if the shared-plus-local approach beats local-only on held-out volunteered ratings. I’d decide the meaningful improvement and minimum evidence threshold before looking at results, not pick them to make a chart look good afterward.

Privacy needs engineering, not a slogan

Model updates can leak information too. A serious pilot would need explicit, revocable opt-in; a minimum group size per round; secure aggregation so the server cannot inspect one Mac’s update; clipping, abuse defenses, and tested differential privacy with a stated budget. If those safeguards aren’t demonstrated, I won’t call it private just because the raw layout stayed put. I’d publish what leaves the Mac and what the server can actually see before asking anyone to join.

That includes the hard cases: too few people in a round, a Mac dropping out halfway through, someone flooding the system with fake clients, and an update designed to poison the model. A usable opt-out must stop future participation and delete local research data. It can’t unteach an aggregate that has already been published, so I’d explain that limitation rather than promise impossible retroactive deletion. Encryption and differential privacy are different protections; neither comes for free just because we call a system federated.

A small pilot might tell us whether people want to participate. It would not justify a claim of a production-grade privacy guarantee or a useful global model. And a placement ranker would not replace the large models that write Mirror, Obituary, or AI Stylist text.

Reasons this might be the wrong idea

Home screens are wildly personal. A shared average might make everyone’s suggestions a little worse, while a tiny model trained just for you might do the job. At small scale, privacy noise can drown out any pattern. Running secure aggregation, accounting for privacy loss, defending against malicious updates, and converting a model for on-device use would cost real engineering time. I won’t use a hypothetical reduction in AI costs as justification today: a placement ranker wouldn’t write the AI text or make the hosted results cap disappear.

So here are my stop signs: not enough people volunteer; explicit ratings are too sparse or biased to evaluate; local-only learning does as well as shared learning; or the privacy protections can’t be demonstrated. Any of those would be a useful result. I’d keep the local feature and leave the federation unbuilt.

What exists right now

Unjiggle does not do federated analytics or federated learning. There is no volunteer training upload. Optional preset ratings stay local. When you explicitly request an AI result, the app can send the documented layout details to Anthropic, either through Unjiggle AI or directly with your own key. That existing AI request is separate from any future research consent. See the privacy policy for the current behavior.

For now, this is the question I’m putting in public: can a shared starting point beat a good local-only one without compromising the trust that made the product worth building? If the answer is no, I’d rather say so and build the simpler thing.

Tell me what sucks about this idea.

If you’ve tried federated learning on personal data—or think local-only is enough—I want to know where this falls apart. What am I missing? Find me @chungty on X.

Challenge me on X →

Opens a draft that tags me and links this post. You decide whether to publish it.