10 Commits
Author SHA1 Message Date
shadowdaoandClaude Opus 5 80b03c884e Use dumb-init --single-child for tini signal parity
Code review caught that dumb-init is not a drop-in replacement for tini as
invoked. tini without -g forwards a signal only to its direct child;
dumb-init without --single-child calls setsid() and forwards to the entire
process group. Verified against dumb-init 1.2.5 with a parent that catches
SIGTERM and does not forward it: under the default the grandchild is still
signalled, under --single-child it is not.

Without the flag, `docker stop` delivered SIGTERM to the tenant's node
process directly and simultaneously with pm2, bypassing pm2's kill_timeout
shutdown sequencing. Since the wedged-container fix in #1 introduced tini
specifically for its signal forwarding, the previous commit's claim that
the two are equivalent did not hold.

Also corrects MEMORY-GUIDE.md, which still described header buffers as
"reduced" after they were raised from 2 1k to the nginx default of 4 8k.
The nginx memory budget is unchanged: these buffers are allocated per
request only when a request needs them, not preallocated per connection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 09:06:10 -07:00
shadowdaoandClaude Opus 5 e417a908f2 Upgrade to AlmaLinux 10 and fix WebSocket proxying
Base image moves from AlmaLinux 9 to 10, which forces two related changes:
EPEL repo URL bumps to the el10 release RPM, and tini is replaced with
dumb-init as PID 1 since tini is not packaged in EPEL 10. Both provide the
signal forwarding and zombie reaping the wedged-container fix relies on.
Nginx 1.26 in AlmaLinux 10 deprecates the `listen ... http2` parameter, so
the server block now uses the separate `http2 on;` directive.

Also fixes three issues that affected WebSocket apps (socket.io, ws):

- `Connection: upgrade` was hardcoded on every proxied request, including
  ordinary HTTP. Now driven by a `map $http_upgrade $connection_upgrade`
  so only genuine upgrade requests carry it.
- No explicit proxy read/send timeout meant idle WebSockets were cut at
  nginx's 60s default. Set to 600s, which clears any sane heartbeat without
  pinning connection slots (worker_connections is 512, two per client).
- `large_client_header_buffers 2 1k` returned 400 for any single header
  line over 1KB, which session cookies and bearer tokens routinely exceed.
  Raised to the nginx default of 4 8k; buffers are allocated on demand, so
  this only costs memory for requests that need it.

Verified by rendering the generated config and running it under nginx with
a Node backend: WebSocket handshakes return 101 and reach the upstream
'upgrade' event, polling requests arrive with `Connection: close`, and a
3KB cookie returns 200 where the old buffer setting returned 400.

README gains a WebSocket Support section covering the PM2 cluster-mode
trap, the concurrency ceiling, and reconnects on memory restarts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 07:54:32 -07:00
shadowdaoandClaude Opus 4.7 b431a66a7b Fix wedged-container outage: TCP healthcheck + tini-managed PID 1
A 10s postgresql restart took down transcribe.shadowdao.com-01 for ~17h
because pm2 gave up after 5 fast retries, the entrypoint's trailing
tail -f kept PID 1 alive, and the healthcheck (wget --spider on nginx
port 80) succeeded on the 301-to-https redirect regardless of whether
Node was alive.

Three coordinated fixes to the cnoc image:

- HEALTHCHECK: replace the redirect-passing wget probe with TCP-level
  checks on 127.0.0.1:3000 (Node) and :80 (nginx). Tenant-agnostic, no
  /ping dependency — catches the exact incident scenario (port 3000
  closed when pm2 exits).
- entrypoint.sh: exec pm2 via tini so it becomes PID 1. When pm2
  exhausts max_restarts and exits, the container exits and the
  unless-stopped restart policy brings it back. Logs are tailed in the
  background with -F (logrotate-safe).
- Dockerfile: install tini from EPEL for proper signal forwarding and
  zombie reaping of nginx/crond children that reparent to PID 1.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 06:59:52 -07:00
shadowdao 566cdb5276 fix generate-ecosystem-config.sh
Cloud Node Container / Build-and-Push (18) (push) Successful in 1m58s
Cloud Node Container / Build-and-Push (20) (push) Successful in 1m56s
Cloud Node Container / Build-and-Push (22) (push) Successful in 1m55s
2025-07-31 12:08:07 -07:00
shadowdaoandClaude a9dd0573ca Fix PM2 UID type error by using login shell and explicit undefined
Cloud Node Container / Build-and-Push (18) (push) Successful in 1m47s
Cloud Node Container / Build-and-Push (20) (push) Successful in 1m52s
Cloud Node Container / Build-and-Push (22) (push) Successful in 2m0s
- Use su with login shell (-) to ensure clean environment
- Explicitly set uid/gid to undefined in ecosystem config
- This prevents PM2 from trying to parse string UID as integer

The error occurred because PM2 was receiving '1002' as a string
instead of an integer. By using a login shell and explicitly
setting uid/gid to undefined, PM2 won't try to switch users.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-24 10:02:17 -07:00
shadowdaoandClaude 68ef106277 Fix PM2 cluster mode and user permission errors
Cloud Node Container / Build-and-Push (18) (push) Successful in 1m52s
Cloud Node Container / Build-and-Push (22) (push) Has been cancelled
Cloud Node Container / Build-and-Push (20) (push) Has been cancelled
- Force PM2 to use fork mode instead of cluster mode
- Disable wait_ready to avoid startup issues
- Add PM2 ready signal to simple-website server
- Add PM2 status check after startup
- Set NODE_ENV=production for PM2 startup

The cluster mode was causing the UID 1002 error. Fork mode runs
the process directly as the specified user without additional
permission complications.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-24 09:53:01 -07:00
shadowdaoandClaude 205c1c4d68 Fix PM2 user permission error (UID 1002 not found)
Cloud Node Container / Build-and-Push (18) (push) Successful in 1m58s
Cloud Node Container / Build-and-Push (20) (push) Successful in 1m51s
Cloud Node Container / Build-and-Push (22) (push) Successful in 1m53s
- Add proper error handling for user creation with detailed logging
- Verify user exists before starting PM2
- Set HOME environment variable when running PM2
- Run PM2 with --no-daemon flag and proper user context
- Add ownership fix for generated ecosystem.config.js

This should resolve the "User identifier does not exist: 1002" error
by ensuring the user is properly created and PM2 runs in the correct context.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-24 09:22:43 -07:00
shadowdaoandClaude 9f0aa4b8b1 Implement auto-generation of ecosystem.config.js and improve container setup
Cloud Node Container / Build-and-Push (18) (push) Failing after 38s
Cloud Node Container / Build-and-Push (20) (push) Failing after 33s
Cloud Node Container / Build-and-Push (22) (push) Failing after 33s
- Add automatic ecosystem.config.js generation from package.json
- Create app directory automatically if missing
- Copy simple-website example when app directory is empty
- Remove redundant default app files from configs/
- Add HAProxy support with proper real IP forwarding
- Configure nginx to trust proxy headers from private networks
- Simplify entrypoint logic - always use /home/$user/app

This makes the container more user-friendly by eliminating the need for
manual PM2 configuration and ensuring the server always has a working app.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-24 09:01:08 -07:00
shadowdaoandClaude e9d4d11177 Fix package conflict by removing curl dependency
Cloud Node Container / Build-and-Push (18) (push) Successful in 49s
Cloud Node Container / Build-and-Push (20) (push) Successful in 54s
Cloud Node Container / Build-and-Push (22) (push) Successful in 1m12s
- Remove curl from package installation to avoid curl/curl-minimal conflict
- Replace curl with wget in health check and Node.js installation scripts
- wget is already installed and provides same functionality

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-21 16:05:37 -07:00
shadowdaoandClaude 2989cd590a Complete Node.js container implementation with multi-version support
Cloud Node Container / Build-and-Push (18) (push) Failing after 10s
Cloud Node Container / Build-and-Push (20) (push) Failing after 8s
Cloud Node Container / Build-and-Push (22) (push) Failing after 8s
- Add Dockerfile with AlmaLinux 9 base, Nginx reverse proxy, and PM2
- Support Node.js versions 18, 20, 22 with automated installation
- Implement memory-optimized configuration (256MB minimum, 512MB recommended)
- Add Memcached session storage for development environments
- Create comprehensive documentation (README, USER-GUIDE, MEMORY-GUIDE, CLAUDE.md)
- Include example applications (simple website and REST API)
- Add Gitea CI/CD pipeline for automated multi-version builds
- Provide local development script with helper utilities
- Implement health monitoring, log rotation, and backup systems

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-07-21 16:00:46 -07:00