My Portfolio Is the Demo
I built my own analytics SDK and a health-data aggregator, then wired both into a single MCP server, so I can ask an AI the questions a dashboard cannot answer.
Most portfolio sites track visitors the same way: bolt on Google Analytics, check the dashboard once a month, feel vaguely informed, move on. I wanted to build the analytics myself, so the portfolio would double as a working example of how this stuff actually runs. And then I wanted what a dashboard cannot do: hand the raw, structured data to a model and let it reason over it.
Then I did it twice. Once for how people use my site, once for my own health data. Both wired into the same AI through one protocol.
The premise
Most side projects exist to show off a skill. This one is the skill, running. Every section you scroll and every nav link you click gets picked up by an analytics SDK I wrote by hand. The site is a working piece of the thing it is describing.
There is a link to the live dashboard in the nav, and that is deliberate. You are a data point right now, and you can go watch yourself show up in it.
The analytics SDK
An analytics SDK has three jobs: figure out who someone is, record what they do, and send it somewhere. Hosted tools make all of those decisions for you and bury them. Building one yourself means you decide each one on purpose.
Two kinds of identity, for two different questions: a session_id is a UUID kept in sessionStorage (t_sid), scoped to a single tab. It answers "is this one person browsing five pages, or five people on one page each?" But a session dies when the tab closes. So there is also a visitor_id, computed server-side as SHA256(ip:user-agent) at an /api/identify endpoint. The client never sees the IP, the server never stores it raw, and the hash is the identity. Hosted tools pick this split for you without telling you, and it quietly shapes every "returning visitor" number downstream.
Events I chose to fire. Fourteen explicit event types cover what I care about: page and section lifecycle, scroll depth, clicks, tab focus, and three specific to the PCA playground. The schema is uniform across all of them:
type TrackerEvent = {
event: EventName
session_id: string
visitor_id: string
page: string
referrer?: string
properties?: Record<string, string | number | boolean>
timestamp: number // client ms epoch
}
I track what I decide matters, not everything that moves. Logging everything does not get you insight, it gets you a pile you pay to query later.
Dwell time, measured properly. Section engagement runs on an IntersectionObserver at a 0.3 threshold, so a section only counts once 30% of it is on screen. On entry I stamp a start time into a Map; on exit I emit the delta. So section_exit carries an actual { section, duration } instead of just "this scrolled into view once." Blog posts get a second layer: per-heading tracking keyed slug#heading-id, so I can see which part of an article holds attention and which part loses it. (That one comes back later.)
Batching, because the network is not free. Scroll fires constantly, so the scroll handler is throttled to 200ms and depth only emits at 25/50/75/100% milestones, once each. Events queue and flush on a 3-second debounce, or when the queue hits 10, whichever comes first. The part that actually matters is unload: a normal fetch gets cancelled when the tab closes, so on pagehide and visibility-hidden I switch to navigator.sendBeacon. Without that you lose the last events of a session, which are usually the ones you care about.
None of this is clever, and that is the reason to build it once by hand. Afterward, "analytics" stops being a black box. It is structured logging with a UI on top.
The backend
Ingest is one POST /api/track. It takes a batch, validates it with Zod, and enriches each event on the server: device, browser and OS from the user-agent, country from Cloudflare's cf-ipcountry header (never from a raw IP I would have to store). Everything lands in one wide Postgres table, portfolio_events, in a single parameterized insert, indexed on the access patterns I actually query: session, visitor, event plus time, page, country.
I do not compute aggregations on the fly. There are nine materialized views: daily stats, session stats, per-page and per-section engagement, blog read-completion, CTA rollups, device and country breakdowns. They refresh CONCURRENTLY from a cron-authenticated /api/analytics/refresh endpoint, with a 60-second Redis cache in front of the dashboard. Geo stats enforce k-anonymity (K_ANON_MIN = 5): a country with fewer than five visitors does not show up. And because watching the numbers move is half the fun, every event is also published to a Redis channel that feeds a Server-Sent Events stream, so the dashboard's "watching now" count is live rather than faked.
What I skipped matters too: no Kafka, no streaming warehouse, no custom query engine. For a personal site's traffic, indexed Postgres plus materialized views answers everything in milliseconds. The point is not "use Postgres," it is to size the system to the traffic you have instead of the traffic you imagine.
Wiring it into an AI with MCP
This is where it stops being an analytics post.
MCP, the Model Context Protocol, is an open standard for giving a model structured, typed access to external tools and data. Instead of pasting a CSV into a chat and hoping, you define tools: functions with typed inputs and outputs the model can call on its own.
I wrapped both backends in one MCP server: Bun, Hono, the official TypeScript SDK, served stateless over HTTP behind a bearer token. Each analytics and health endpoint becomes a tool. The detail that makes it work is the output schema:
server.registerTool(
'analytics_engagement',
{
description:
'Section dwell times, session-depth histograms, and home→blog/projects conversion funnels.',
inputSchema: { days: z.number().int().min(1).max(365).optional() },
outputSchema: {
sections: z.array(sectionStat), // section, enter_count, avg/median/max dwell
session_summary: sessionSummary, // medians, means, duration + page histograms
funnels: z.array(funnelStat), // home_to_blog, home_to_projects
},
annotations: { title: 'Engagement Analytics', readOnlyHint: true },
},
({ days }) => run('analytics_engagement', () => client.getDashboard(days)),
)
The typed outputSchema is what lets the model chain tools without me holding its hand. It knows analytics_engagement returns funnels before it calls it, so it can plan ahead: pull engagement, spot a weak funnel, then call analytics_blog to dig in, all in one turn without being told to. (One thing that bit me: I had to loosen a couple of enums to z.string(). The moment I shipped a new funnel ID, the strict schema rejected the entire payload. Typed contracts help right up until the upstream data shifts under them.)
The change from a dashboard is small to describe and large in practice: the model decides what to look at. I do not open a chart or remember which metric lives where. I ask a question and it fetches what it needs.
Letting the model reason
With the server connected, I asked Claude something open-ended: "give me interesting insights on user behavior." It pulled engagement, CTA, and blog data in one turn and pointed at a few things I had scrolled past for weeks. (The site's traffic is still small, so read the specifics below as directional, not gospel.)
Mean and median tell different stories. Take session length. The median can sit down in the seconds while the mean runs up into the minutes, which only happens when a long tail of very engaged sessions drags the average way up. Most people bounce fast; a small group stays a long time. A single "average session" number hides exactly the group you care about. The model read that off the distribution straight away, where a bar chart would have flattened it.
The "Analytics" nav link out-clicks "Home." The kind of person who finds this site is more curious about the meta-layer than about navigating it. That points at what belongs up front.
Projects convert better than blog posts off the homepage. A visitor landing on the homepage is noticeably more likely to click into a project than into a post. The data says there is a gap, not why. Weak thumbnails, buried section, flat copy. It points the flashlight; the digging is on me.
The PCA post loses readers at the math. The per-heading tracking pays off here: on a formula-heavy post, attention falls off right where the math starts. That is measurable, and it tells me what to do: add intuition before the formulas, or split the post in two.
None of this needed a custom insight engine or a line of analysis code. It needed clean, structured access and one decent question.
Crossing domains
The same MCP server also wraps a health-data aggregator I built. An iPhone Shortcut pushes Apple HealthKit data to a POST /health/sync endpoint: resting heart rate, VO2 max, step count, and full running workouts with per-kilometer splits and GPX tracks. The endpoint is idempotent by design. It is an upsert keyed on (metric_type, metric_date) that skips the write if the value has not changed, so re-syncing a month costs nothing. Workouts get an effort score from the TRIMP (Training Impulse) formula -- duration × avgHR × 0.64 × e^(1.92 × HRR), where HRR is heart rate reserve -- which rolls up into acute and chronic training load (ATL/CTL), the ratio coaches watch for overtraining. Same stack as the analytics backend: Bun, Hono, Postgres, materialized views, Railway, fail-open Redis.
So now the model has tools across both domains, analytics_engagement sitting next to health_get_week_summary, behind one bearer token. The question I actually wanted to ask: do my high-output writing days line up with my high-activity days, or do I trade one for the other?
No single tool answers that. It is a join across two datasets that were never meant to meet, done by a model that can hold both at once. That is the actual point of the whole build.
Why this is the right shape
Dashboards are great for questions you already know to ask. What is my traffic today? How many US visitors? Those are lookups, and charts win at lookups.
The questions worth having are the ones you do not know to ask yet. Why do my long sessions cluster on certain entry pages? Does a heavy training day hurt my writing the next morning? Those are not lookups, they are reasoning problems, and models are good at reasoning over structured data once they have the tools to go fetch it.
MCP is what keeps this composable. I add an endpoint, wrap it as a tool with a typed schema, and the model can use it right away, with no new prompt and no new dashboard. It decides when the tool is relevant. The part I am proud of is the architecture, not any single insight it turned up.
What's next
Two directions. First, more history on both sides, enough that the model can pick up seasonal patterns and not just day-to-day ones. Second, a read-only, rate-limited slice of the analytics MCP exposed publicly, so other developers can point their own Claude at this site's live data and try it for themselves.
The portfolio started as a place to park my projects. It ended up being the one I tinker with most.
The analytics SDK, both backends, and the MCP server are all hand-built. The nav link up top shows the analytics live, so you are in the data already. Ask me about any piece and I will write the follow-up.