This document explains how Beyou defends itself: how users prove who they are, how every request is validated and throttled, how destructive actions demand a second factor, and which guards refuse to even start the server when production is misconfigured. It ends with an honest assessment of what is still missing.
One framing note first: the backend never terminates TLS. HTTPS, and therefore the safety of every cookie and header below, is the reverse proxy's job in front of the loopback-bound containers. The infrastructure topic covers that layer.
flowchart LR
subgraph client["Client"]
FE["⚛️ Web / 📱 Mobile<br/>JWT in memory"]
end
subgraph filters["Request pipeline"]
RL["🚦 RateLimitFilter<br/>bucket4j tiers"]
SF["🛡️ SecurityFilter<br/>JWT validation"]
DS["🔏 DocsImportSecretFilter"]
end
subgraph server["Server side"]
LA["🔒 Login lockout<br/>per account"]
TS["🔑 TokenService<br/>HMAC256, 15 min"]
RT["🔄 Refresh tokens<br/>BCrypt-hashed, rotated"]
OWN["👤 Ownership checks<br/>in every service"]
end
FE -->|"Authorization: Bearer"| SF
FE -->|"refresh cookie / body"| RT
SF --> RL --> OWN
SF --> TS
FE <-->|"OAuth 2.0"| GO["🔐 Google"]
LA -.-|"guards"| TS
DS -.-|"guards /docs/admin"| OWN
The design in five lines:
X-Access-Token response header.| Endpoint | Method | Auth | Purpose |
|---|---|---|---|
| /auth/login | POST | No | Email + password login |
| /auth/register | POST | No | Registration (email verification required before login) |
| /auth/verify-email | GET | No | Consume the 24-hour verification token |
| /auth/resend-verification | POST | No | Issue a new verification token and mail it; always the same 200 |
| /auth/google | GET | No | Google OAuth code exchange (web) |
| /auth/google/mobile | POST | No | Google ID-token verification (mobile) |
| /auth/refresh | POST | No | Rotate the refresh token, mint a new JWT |
| /auth/logout | POST | No | Clear the cookie, revoke the token |
| /auth/verify | GET | Yes | Session probe; returns "authenticated" |
| /auth/forgot-password | POST | No | Request a reset e-mail |
| /auth/reset-password/validate | GET | No | Pre-check a reset token |
| /auth/reset-password | POST | No | Set the new password |
Every unauthenticated auth endpoint shares one rate bucket: 5 requests per 15 minutes per client IP.
sequenceDiagram
participant U as User
participant BE as Backend
participant DB as Database
U->>BE: POST /auth/login
BE->>BE: Lockout check (10 fails / 15 min per email)
BE->>DB: Find user by email
BE->>BE: BCrypt.matches(input, hash)
BE->>BE: emailVerified? else 403 EMAIL_NOT_VERIFIED
BE->>DB: Create refresh token (hash stored)
BE-->>U: JWT in X-Access-Token + refresh cookie
The order of the checks is the interesting part:
matches always fails.Registration stores the user with a 32-byte verification token (24-hour expiry) and sends the confirmation e-mail. The token is single-use: consuming it sets emailVerified and nulls both token columns. Password policy is enforced in the service layer, with the DTOs as backup: at least 12 characters and at least 2 of the 4 character classes.
An honest tradeoff, stated in the code: registration answers "Email already in use" for a taken address, so it is an enumeration oracle by design. The 5-per-15-minutes IP bucket is what keeps that from being farmable at scale.
Two separate paths, one per platform:
GET /auth/google?code=): the backend exchanges the authorization code with Google server-side (client secret never leaves the server) and reads the profile with the resulting access token. The web client generates and verifies its own state value before handing the code over.POST /auth/google/mobile): the native app sends a Google ID token, which the backend verifies with Google's official verifier: signature against Google's published keys, issuer, expiry, and an audience allowlist. The token is additionally rejected unless Google itself reports the e-mail as verified.Both paths find-or-create the user by e-mail. Google-created accounts get isGoogleAccount=true, a non-hash password marker, and skip e-mail verification (Google already did it).
Both paths also refuse a matched account that is a password account with an unverified address, returning the same 403 EMAIL_NOT_VERIFIED that login does. Until that guard existed, doLogin was the only reader of emailVerified in the backend, so Google was a way around the gate — mildly, as an accidental cure for a lost verification mail, and seriously as this: anyone can register an address they do not own, and the unverified row they leave behind would swallow the real owner's Google sign-in with no click and no warning. The owner fills that row with their data, and if they ever follow the verification link that arrived when the squatter registered, the flag flips and the squatter's password opens the account. One rule now holds at every door, and it is recoverable rather than merely strict because the resend endpoint landed with it. A verified password account may still link Google freely.
| Property | Value |
|---|---|
| Algorithm | HMAC256 (auth0 java-jwt) |
| TTL | 15 minutes |
| Claims | iss=auth-api, sub=email, exp. Nothing else |
| Delivery | X-Access-Token response header |
| Consumption | Authorization: Bearer request header |
| Storage | Frontend memory only |
The claims are deliberately minimal. The role is not in the token; the SecurityFilter re-reads the user row on every request, so a role change or a deleted account takes effect within one request rather than one token lifetime. The cost is a database read per authenticated request, which the Caffeine layer absorbs elsewhere but is a real trade here.
HMAC256 over RSA because only this backend ever signs or verifies: there is no third party to hand a public key to.
The client-held token is {rowId}.{secret}: a UUID naming the database row plus 32 random bytes. The database keeps only the BCrypt hash of the secret, so a leaked table contains nothing replayable.
flowchart TD
CR["🔑 32 random bytes"] --> HASH["🔒 BCrypt hash (cost 12)"]
HASH --> DB["💾 Row: id + hash + expiresAt + revokedAt"]
CR --> OUT["📤 To client: id.secret"]
OUT --> REF["🔄 POST /auth/refresh"]
REF --> MATCH["matches(secret, hash)?<br/>expired? revoked?"]
MATCH --> ROT["Revoke old row, mint new pair"]
Secure and SameSite=Strict in production (Lax in dev), path /, 15-day maxAge.X-Client: mobile, the backend skips the cookie entirely and returns the refresh token in the response body; later refreshes send it back in an X-Refresh-Token header. Cookies are a poor fit for native HTTP stacks, so mobile owns its storage.Three custom filters cooperate, and their order matters:
| Order | Filter | Job |
|---|---|---|
| 1 | SecurityFilter (before UsernamePasswordAuthenticationFilter) | Bypass list for public paths; otherwise extract the Bearer token, validate signature/expiry/issuer, load the user, populate the SecurityContext. Failures answer 401 with a keyed ApiErrorResponse (JWT_NOT_FOUND, AUTH_HEADER_INVALID, JWT_INVALID, USER_NOT_FOUND) |
| 2 | DocsImportSecretFilter (after UsernamePasswordAuthenticationFilter) | Constant-time comparison of the X-Docs-Import-Secret header for /docs/admin/import/*; a blank configured secret fails closed with 403 |
| 3 | RateLimitFilter (plain servlet filter) | Runs after the security chain, which is exactly what lets it key buckets by authenticated user |
Two details worth knowing before touching this code. First, the public-path list exists twice: once as permitAll matchers in SecurityConfig and once as the SecurityFilter's bypass conditions. They agree today, but they match differently (equals vs startsWith), and drift between them is silent. Second, async dispatches are permitted through the chain because the agent's SSE stream re-dispatches; the compensating invariant is that every protected endpoint must authenticate and ownership-check on the initial dispatch.
Bucket4j buckets in a Caffeine cache, first matching tier wins:
| Tier | Endpoints | Limit | Keyed by |
|---|---|---|---|
| auth | login, register, forgot-password, resend-verification, google, google/mobile | 5 / 15 min | IP |
| agent | POST /ai/agent/chats/* | 30 / hour | user |
| docs | /docs/* (public) | 30 / min | IP |
| photo | GET /user/photo/* | 120 / min | IP |
| onboarding | POST /onboarding/suggestions | 30 / hour | user |
| account-deletion | POST /user/deletion/* | 10 / hour | user |
| feedback | POST /feedback | 10 / hour | user |
| feedback-attachment | POST /feedback/*/attachments | 20 / hour | user |
| export | GET /user/export | 5 / hour | user |
| write | any other POST/PUT/DELETE | 30 / min | user |
| read | any other GET | 60 / min | user |
The export sits above the generic read tier for a reason worth stating: it is a GET, but it returns the entire account in one response — every category, habit, task, goal, routine, feedback thread and assistant conversation, assembled in memory and serialized in one go. Sixty a minute of that is a way to hold the heap, and nobody taking their data needs a sixth copy inside the hour.
Rejections answer 429 with a Retry-After header; successes carry X-Rate-Limit-Remaining. Both are named in Access-Control-Expose-Headers, without which a browser cannot read either one: neither is on the CORS safelist, so the wait was on the wire and unreachable by the web client.
The client IP comes from the CF-Connecting-IP header, not X-Forwarded-For, and the reason is worth remembering: Cloudflare appends to X-Forwarded-For rather than replacing it, so its leftmost entry is attacker-controlled, and honoring it would hand out a fresh login bucket per request. When the header is absent the filter falls back to the socket address, which behind a tunnel collapses into one shared bucket. That degraded case is precisely why the per-account login lockout exists as an independent second layer.
The whole subsystem is off in the e2e and test profiles, and user-keyed tiers pass unauthenticated requests through untouched.
{rowId}.{secret}, BCrypt-hashed at rest, single-use, 15-minute TTL, and requesting a new one invalidates all previous tokens.Deleting an account is the one action where a logged-in session is deliberately not enough: the flow demands proof of inbox access.
POST /user/deletion/code mails a six-digit code. BCrypt-hashed at rest, 15-minute TTL, 60-second cooldown between requests, and each new code invalidates the previous ones.POST /user/deletion/confirm checks, in order: already used, expired, too many attempts (5), then the hash comparison. The attempts counter increments in its own REQUIRES_NEW transaction, because the exception that follows a wrong guess rolls the outer transaction back, and counting inline would have left the cap unreachable.There is no method-level security in the codebase, on purpose. The model is one rule applied everywhere: every service method receives the authenticated user's id and compares it against the loaded entity's owner, throwing a keyed error on mismatch (CATEGORY_NOT_OWNED, HABIT_NOT_OWNED, TASK_NOT_OWNED, GOAL_NOT_OWNED, ROUTINE_NOT_OWNED, SNAPSHOT_NOT_OWNED, CHAT_NOT_OWNED, FEEDBACK_NOT_OWNED). Schedules route through the owning routine, which is what closed an early IDOR. These all surface as HTTP 400 with an errorKey; clients discriminate on the key, not the status.
Exactly one role rule exists: /feedback/admin/** requires ADMIN. The ADMIN role is granted only by a manual database update. No seed, no endpoint, no environment variable can mint an admin.
The two upload paths (profile photo, feedback attachments) share the same defensive shape:
A photo is stored in two unrelated places and read in priority order, and that is the whole reason removal needed its own endpoint. An upload writes {upload-dir}/user-photos/{userId}.jpg and never touches the user row; perfilPhoto on the row holds a Google CDN URL, set only at OAuth sign-in. UserMapper looks for the file first and falls back to the column.
DELETE /user/photo clears both. Removing one half always leaves a photo on screen: drop only the file and a Google account falls back to the avatar it had before, clear only the column and the uploaded file goes on being served. The second case is also why PUT /user with an empty photo never worked as a removal, which is what users hit.
The file is unlinked before the column is cleared, and a failed unlink rolls the whole thing back. The alternative order can commit "this account has no photo" over a JPEG that is still on disk and still winning the priority check, which is the one outcome worse than refusing.
The account id comes from the token, never from the path, so the endpoint has nothing of the enumeration surface GET had to be signed to close.
Reading a photo back is the one place here where authorization does not travel in a header. The callers are an <img src> on the web and an <Image uri> on the phone, and neither can send one, so GET /user/photo/{userId} used to answer any caller who could name a user id. Every uploaded face was readable by walking the UUID space.
The URL carries its own proof instead:
/api/v1/user/photo/{userId}?v={mtime}&exp={epoch}&sig={HMAC-SHA256(userId|exp)}
HMAC(TOKEN_SECRET, "beyou-photo-url-v1"), so there is no second secret to deploy, and a photo signature is useless as a token anywhere else.UserMapper mints the URL while answering GET /user, and nothing else mints one. Login does not: it maps the user without a photo version, so a client that wants the signed URL has to ask for the profile.exp is covered by the signature, so the deadline cannot be extended by editing the query string. The default TTL is 12 hours (PHOTO_URL_TTL_MINUTES), which keeps an avatar rendering in a tab left open overnight while a URL captured in a proxy log stops working the same day.MessageDigest.isEqual, so a partial guess leaks nothing about how much of it was right.Cache-Control is private, because a shared cache would go on serving the bytes after the signature expired.The cost is that the URL works for whoever holds it until exp passes, including anyone it gets forwarded to. It exposes a single image the sender could already see.
The agent chat can call real tools, so its authority model matters:
Headers set by the backend on every response:
default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; img-src 'self' https: data:; connect-src 'self' https://accounts.google.com https://www.googleapis.com; font-src 'self' https: data:; frame-ancestors 'none'CORS: one allowed origin pattern from the environment, credentials enabled, and exactly one exposed header: X-Access-Token. Dev runs a wildcard; production refuses one (next paragraph).
Boot-time validators, the "refuse to start" layer:
| Guard | Refuses boot when |
|---|---|
| SecurityConfigValidator (prod only) | CORS pattern is *, JWT secret shorter than 32 chars, cookie.secure false, or either e2e escape hatch (deletion-code exposure, auto-verified e-mail) is enabled |
| SchemaOwnershipGuard | Flyway is on but Hibernate ddl-auto is anything other than validate or none |
| E2eSafetyCheck (e2e profile) | The datasource URL does not look like a test database |
Operational posture: the actuator lives on its own loopback-bound port with a fixed endpoint list in production (the env override is deliberately dropped there); Swagger is off in production; AOP logging records argument counts, never values; the container runs as a non-root user; CI runs CodeQL and a weekly OWASP dependency check.
| Area | Current state | Honest note |
|---|---|---|
| 2FA / MFA | Not implemented | The deletion flow's e-mail code is the only second factor in the product |
| Audit logging | Not implemented | Failed logins, resets, and refreshes leave no dedicated trail |
| Refresh token binding | Not bound to device or IP | Rotation limits the damage window but a stolen token works anywhere until then |
| Google account linking | Find-or-create by e-mail, verified accounts only | A VERIFIED password account is still logged in by a matching Google identity with no explicit linking step. The unverified case, which was the dangerous one, is now refused |
| Registration enumeration | "Email already in use" by design | Rate-limited, and a usability tradeoff, but still an oracle |
| Reset cooldown nuance | 400 inside the cooldown for real accounts | A patient prober can distinguish known addresses on a second request |
| verify-email throttling | Unthrottled | Unauthenticated GET that falls through the user-keyed tiers; token entropy is the only guard. Its sibling POST /auth/resend-verification IS in the auth tier |
| Verification token at rest | Plaintext column on the users row | The reset token is stored as a BCrypt hash; this one is readable straight out of a database dump |
| Docs import secret | Compared constant-time, fails closed when blank | Nothing validates its length or entropy at boot |
| Prompt injection | Instruction-level defense only | No programmatic filtering of user text before it reaches the model |
| CSP regression test | Header existence is asserted, value is not | A silent CSP weakening would pass the suite |
| Threat | Mitigated? | How |
|---|---|---|
| Password theft from a DB leak | Yes | BCrypt cost 12; refresh/reset/deletion secrets stored as hashes too |
| XSS stealing tokens | Mostly | JWT in memory, refresh in HttpOnly cookie, CSP on API responses |
| CSRF | Yes | Stateless bearer auth; cookie read only by refresh/logout; SameSite Strict in prod |
| Brute-force login | Yes | 5/15min per IP plus 10-failure account lockout |
| User enumeration | Mostly | Login and reset are silent; registration and the reset cooldown are the documented exceptions |
| Token replay after rotation | Yes | Old refresh tokens are revoked transactionally |
| IDOR | Yes | Ownership check in every service, keyed errors, schedule routed through its routine |
| Decompression bombs | Yes | Header-level pixel cap before decode |
| Confused-deputy AI tools | Yes | Server-built ToolContext; tools inherit caller authority only |
| Session fixation | Yes | No sessions exist |