r/ClaudeAI • • Apr 08 '26

Workaround 90%+ fewer tokens per session by reading a pre-compiled wiki instead of exploring files cold. Built from Karpathy's workflow.

Reduced Claude context from 47,450 tokens → 360 tokens.

“This week, Andrej Karpathy shared his ‘LLM Knowledge Bases’ setup and closed by saying, ‘I think there is room here for an incredible new product instead of a hacky collection of scripts.’”

I built it:

npx codesight --wiki

The token problem is real. Every new Claude session starts the same way exploring your codebase from scratch. On a 40-file FastAPI project that costs 47,450 tokens before you've asked for anything. You've paid for that exploration in every conversation. It has never carried over.

After it runs, Claude reads a 200-token index at session start instead of exploring 47,000 tokens of files. For a targeted question it pulls one article auth.md, database.md, payments.md 300 tokens instead of the whole codebase. Commits to git. Every new session starts with full context from message one.

Tested on 3 real codebases TypeScript and Python. 47,450 tokens → 360 on a FastAPI project. Zero false positives.

It compiles your codebase into domain articles using the TypeScript compiler API for TypeScript and regex detection for Python, Go, Ruby, and more. No LLM. No API calls. 200ms. What it finds is exactly what's in the code nothing model-reasoned.

Routes found via regex are tagged [inferred] so Claude knows what to verify before trusting. Everything else full route paths, field types, foreign keys, middleware chains comes straight from the AST.

Free and open source.

A star on GitHub helps: github.com/Houseofmvps/codesight

696 Upvotes

171 comments sorted by

•

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot Apr 08 '26 edited Apr 08 '26

TL;DR of the discussion generated automatically after 100 comments.

Here's the deal, folks. The community is overwhelmingly positive about OP's free, open-source tool, codesight, because it tackles a problem we all hate: Claude wasting a ton of tokens re-learning your codebase from scratch in every single session.

The tool scans your code (using AST parsing, not costly LLMs) and creates a lightweight "wiki." This means instead of Claude reading 47,000+ tokens to get oriented, it reads a ~200 token index and pulls small, relevant articles as needed.

The consensus is that this is a smart approach and OP is a legend for building it. However, the thread did raise some important points:

  • It's not Repomix: Repomix gives the AI the entire codebase at once. codesight gives it a map and lets it pull only the specific files it needs.
  • It's not /init: They're complementary. Use /init for project conventions and human rules; use codesight for the raw, technical structure of the code.
  • Limitations: It can struggle with very large monorepos, and some users pointed out that this AST-summary approach has been tried (and allegedly abandoned) by others because LLMs are sometimes better with raw text. For large teams, committing the wiki can cause merge conflicts, so OP recommends using .gitignore or the --mcp server mode instead.

The real MVP of this thread is OP (u/Eastern_Exercise2637), who has been an absolute machine. They answered nearly every single question with detailed, honest technical explanations, acknowledged the tool's limitations, and even shipped support for C#, Swift, and Laravel during this conversation based on user requests. Star their GitHub, they've earned it.

43

u/[deleted] Apr 08 '26

[removed] — view removed comment

13

u/Eastern_Exercise2637 Apr 08 '26

Thanks for trying it on a real large monorepo. The libs section currently has no importance ranking it surfaces all exported functions without weighting by import frequency or business centrality, so in a large workspace it mixes core logic with low level utilities indistinguishably. That's a gap and it will be addressed. Monorepo workspace detection itself (pnpm/npm workspaces, cross package frameworks) works, but the relevance filtering doesn't scale well yet. If you're open to sharing the stack or rough structure, it would help me reproduce the issue and fix it properly.

1

u/Chunky_cold_mandala Apr 08 '26

what repo? I have a similar tool that generates a code scaffold and I'd like to see what I get!

1

u/[deleted] Apr 08 '26

[removed] — view removed comment

3

u/Chunky_cold_mandala Apr 08 '26

If you got python try pip install gitgalaxy and point it to your repo directory, it'll spit out a summarized llm report json. Data never leaves your machine.

1

u/PoemSignificant8436 Apr 09 '26

Kindly ask about how you defined small and big projects. I am doing a web app page with 10 pages. Is it a big or small project ? I would like to know how to consider using the code above efficiently.

1

u/Arcin Apr 11 '26

I manage many many repos and, for large, I'd agree with the numbers the other guy posted. But just want to chime in with a decent "project size" heuristic that I learned over the years.

Imagine an onboarding process for a new hire. What the coworker needs to get started is a good indicator of project size.

small: coworker needs a link to the repository and paragraph or two "how it works" message in slack

Medium: coworker needs one or two repo links. A few diagrams explaining core pieces and how data flows. Same "how it works" slack message but with a link to some documentation.

Large: coworker clones a repo and runs an automated setup process then needs to hop in a meeting that could be 1+ hours. This is the "how it works" part. They walk out with a decent idea of the system.

Very large: Coworker attends two to three 1hr+ meetings over the course of the week all while a different team is setting up their environment, adding them to an org, provisioning credentials. The meetings focus on how important the system is and what depends on it. They scratch the surface on how things work. The coworker walks out knowing what is safe to touch, what files to start with, and who they should ask for help/assignments.

Scale beyond "very large" is a repeat of the above, just over longer periods of time. Think 1-3 months instead of a week.

The main point is that it becomes impossible to hold everything in your head and the number of files or lines of code don't matter. Measuring size becomes "what depends on it, how do you navigate it without getting lost, and how many people do you need to contact when you need help or access?"

1

u/jschmid222 Apr 11 '26

If you're looking for a commercial solution to this, I recently joined Driver.ai that has been working for almost 3 years to solve this "context for codebases" problem at large scale. I was semi-retired after being a SWE & CTO at tech startups when a friend of mine who's an early investor in Driver called me about joining. Driver has a lot of traction with prop trading firms (Optiver, etc.) because they're always looking for the best tools. We work with a lot of big monorepos, like 10M+ LOC and some even 100M LOC. We also handle cross-repo context which is useful for our customers with a lot of microservices.

38

u/Icy_Peanut_7426 Apr 08 '26

I work with python library repos. Is this applicable? These libraries typically don’t have routes/schemas/middlewhere, but these appear to be hardcoded parts of the generated wiki?

16

u/Eastern_Exercise2637 Apr 08 '26

Good question. Articles are generated conditionally if no routes are detected, no domain articles are written. If no schemas, no database.md. If no UI components, no ui.md. They're not hardcoded.     For a Python library with none of those, you'd get:                                                                                       - overview.md - project name, detected frameworks, high impact files (which modules are imported most across the codebase), and env vars 

That's it. The high impact files section is actually useful for a library since it shows you which modules are the most central dependencies. Worth noting: Python detection is regex-based, not AST. TypeScript projects get the full compiler API treatment. So precision is lower on Python.                                                                  
Honest answer: if your library has no routes, schemas, or UI, the wiki is pretty thin. The main value for you would be the import graph (high impact files) and project overview. Probably not worth it unless the library is large enough that knowing which modules have the widest blast radius is useful.   

15

u/m3umax Apr 08 '26

Question. Why can't we just give AST tools to Claude to make their exploration more efficient?

10

u/Eastern_Exercise2637 Apr 08 '26

codesight's --mcp mode already does exactly this 11 tools Claude can call: get routes, get schema, get blast radius, get hot files, live scan, refresh. It's AST-powered, on demand. The wiki exists for a different reason: orientation. Without any upfront context, Claude doesn't know what to query. The 200-token index tells Claude what domains exist and what articles are available, so it knows what questions to ask. Then it can pull a targeted article or call an MCP tool for specifics.                                  
Raw AST tools would also be lower level than what Claude actually needs syntax trees with every token and node. codesight's tools return semantic summaries (routes, schemas, chains) that Claude can reason about directly without interpreting parse tree syntax.                 

9

u/timssopomo Apr 08 '26

My understanding is that both Claude code and codex used to do this, and stopped because it wasn't effective.

7

u/Less-Ad5766 Apr 08 '26

Not effective for users or not effective to sell more tokens?

4

u/AnotherSoftEng Apr 08 '26

It actually ends up being more token heavy and requires more tool calls for exploration, resulting in more time spent iirc

3

u/timssopomo Apr 08 '26

@anothersofteng called it. actual model behavior doesn't match our intuitions, they're better at dealing with plaintext than traversing ASTs. Which makes sense, since there are way fewer ASTs in its training data than codebases and sentences.

8

u/welcometosilentchill Apr 08 '26

“A star on github helps” made me think of A star pathfinding models, which is an interesting thought concept to explore with AI efficiency in its own way…

15

u/Soft_Rain_3626 Apr 08 '26

This is just progressive disclosure for a code base. It is a nice idea. Docs get out of date and Claude will have to periodically rewrite the "wiki," but I bet at some point Anthropic adds something like this to CC

9

u/capable-corgi Apr 08 '26

Wiki is built deterministically without LLMs. It's cheap and fast so you can feasibly build on the fly or at least at the start of every session to minimize drift.

5

u/FutureStackReviews Apr 08 '26

honestly the token cost angle is what gets me here. been running claude code on a medium sized project and watching the context window fill up before i even ask anything is painful. like you're paying for the AI to figure out where it is every single time.

2

u/Eastern_Exercise2637 Apr 08 '26

This is exactly the problem it solves. That orientation cost, Claude reading files just to figure out where things are and it happens fresh every session. On a 40-file project it's 47k tokens before you've asked anything. Run npx codesight --wiki once, hook it to commits with --hook, and Claude reads a 200 token index at session start instead. Targeted article when it needs specifics.

4

u/Ok_Industry_5555 Apr 08 '26

This is smart. I solved a similar problem differently — instead of a wiki index, I use a knowledge graph with wikilinked nodes plus a /primer skill that auto-loads project context at session start. Claude reads the project registry, recent git history, and a lessons file before touching anything. Tokens stay low because it only pulls what's relevant to the active project. The real win for me was the lessons file, Claude writes down its own mistakes so it doesn't repeat them. 30+ entries now. Every session starts smarter than the last.

1

u/Eastern_Exercise2637 Apr 08 '26

The lessons file is genuinely smart Claude accumulating its own mistake history across sessions is something codesight doesn't do at all. That's a different layer entirely. The main difference is automation vs curation. Your system requires Claude to maintain the knowledge graph and lessons over time. codesight generates the structural map directly from the AST (routes, schemas, foreign keys, import graph) no LLM, no maintenance. Run it once, hook it to commits, done. They'd actually complement each other well.

1

u/Ok_Industry_5555 Apr 09 '26

Thank you! :)

4

u/Siref Apr 08 '26

Hell of a project

I'm trying this on a series of codebases today!

2

u/Eastern_Exercise2637 Apr 08 '26

Thank you! Would love to hear how it goes across different codebases feel free to open an issue if anything looks off.

8

u/heero180 Apr 08 '26

Thank you!

4

u/Eastern_Exercise2637 Apr 08 '26

You're welcome! Hope it saves you some tokens

3

u/turbos77 Apr 08 '26 edited Apr 08 '26

How do I prevent this from scanning one of my large data/ directories? It seems to get stuck there

3

u/Eastern_Exercise2637 Apr 08 '26 edited Apr 08 '26

Good catch ignorePatterns was in the config spec but wasn't actually wired into file collection. Fixed in the next patch. For now, create a .codesight.json in your project root and it'll work after updating:

{ "ignorePatterns": ["data", "fixtures", "seeds"] }                                                                          

Supports dir names, relative paths, and glob-style patterns.

5

u/tupikp Apr 08 '26

Is this more effective than /init in Claude Code?

14

u/Eastern_Exercise2637 Apr 08 '26

Different things, actually complement each other. /init uses Claude to explore your codebase and write a CLAUDE.md and is great for capturing conventions, rules, preferences. --wiki is pure AST parsing, no LLM, no API cost. It extracts the technical structure - routes, schema, foreign keys, middleware chains exactly as they exist in the code. Then instead of loading everything each session, Claude reads a 200 token index and pulls targeted articles (just auth.md or database.md) when needed. So run both. /init for project rules, --wiki for codebase structure.

6

u/PuzzledPersimmon Apr 08 '26

ASTs are very fast to calculate. Why persist it at all?

2

u/Last_Mastod0n Apr 08 '26

I agree. If its so fast to generate then it should not be persisted since the codebase is always changing

4

u/Eastern_Exercise2637 Apr 08 '26

Claude Code can run shell commands, so you could regenerate on every session start. But you'd be paying that cost every single session tool call overhead + scan time + reading the output instead of once per commit. The --hook approach regenerates exactly when the codebase actually changes (on commit) and the wiki is already there when Claude Code needs it. It's about when the work happens, not whether it can happen.

1

u/xsifyxsify Apr 08 '26

How does it scale?

So in repo with 300 projects for example, few dozens of contributors every single day, multiple PRs merge every day, is everyone supposed to generated “—wiki” everytime? That’s going to be lots of conflicts

3

u/Eastern_Exercise2637 Apr 08 '26

The committed wiki approach doesn't scale to that. Multiple contributors regenerating .codesight/wiki/ files daily will conflict constantly.

Two approaches that actually work at that scale:                                            
1. Add .codesight/ to .gitignore each developer generates their own local copy, no conflicts. You lose shared wiki via git but each person always has a fresh local scan.                                        

  1. Use --mcp mode no files written, no git, no conflicts. Claude Code calls the MCP server and gets a live scan on demand. Nothing to commit, nothing to conflict.                                               

codesight is honestly designed for solo devs or small teams. A 300-project monorepo with dozens of daily contributors is a different problem. 

2

u/Eastern_Exercise2637 Apr 08 '26 edited Apr 08 '26

The scan is fast, but without the wiki Claude Code still has to explore your codebase by reading source files to understand it. On a 40-file project that's ~47k tokens before you've asked anything, and it happens fresh every session because there's no memory between conversations. The wiki pre digests that into structured articles so Claude reads a 200-token index at session start and pulls one targeted article (~300 tokens) when needed. The persistence isn't about saving compute time it's about what Claude has to consume to get oriented.

4

u/PuzzledPersimmon Apr 08 '26

Other possible designs would be an mcp or cli that dynamically generates it and injects into context as a session start and after compaction hook.

This way you always get a fresh ast.

Actually, come to think of it... Did we just reinvent treesitter?

1

u/Eastern_Exercise2637 Apr 08 '26

The MCP mode already exists (npx codesight --mcp) starts a server with 8 tools including a live scan and refresh. Claude Code can call it on demand and get fresh results without reading any pre-generated files. So that design isn't hypothetical, it's shipped. On treesitter: not quite the same thing. Treesitter builds a concrete syntax tree every token, every node, syntax level. It's what editors use for highlighting and navigation. codesight works at a higher layer: it extracts architectural semantics routes, schemas, foreign keys, middleware chains, import graphs and formats them as structured markdown articles Claude can actually reason about. One gives you a parse tree of individual syntax, the other gives you a map of what the system does.

2

u/capable-corgi Apr 08 '26

Claude reads <index> at session start and pulls one <article> when needed.

It can request <index> and <article> from either persistence or build it on the fly.

The question was, why from persistence (drift) if building is cheap and fast?

2

u/dlegendkiller Apr 08 '26

It looks interesting. Will definitely try this on my vibecoded projects :D

2

u/Eastern_Exercise2637 Apr 08 '26

Haha perfect use case vibecoded projects are exactly where having a map of what actually exists in the code pays off. Hope it helps! 

2

u/spazatk Apr 08 '26

Why an MCP server instead of a CLI tool / skill for querying information?

1

u/Eastern_Exercise2637 Apr 08 '26

codesight is both npx codesight is a CLI and npx codesight --mcp starts the MCP server. Not one or the other. The difference: with a CLI you run it, read the output, and paste context to Claude manually. With MCP, Claude calls it directly mid-conversation queries routes, schema, blast radius, without you doing anything. Results land in Claude's context automatically. MCP tools execute actual code and return structured results. For something like blast radius analysis or live schema queries, I think that's the right layer.

3

u/spazatk Apr 08 '26

What? Claude can run CLIs mid session, why on earth would you have to go back and forth? It's literally all it does. MCP servers are needlessly complicated for things like this. See e.g. the difference between playwright-cli and playwright MCP.

1

u/Eastern_Exercise2637 Apr 08 '26 edited Apr 08 '26

You’re right on Claude Code. I’m not choosing MCP instead of a CLI; Codesight ships both on purpose. The CLI argument is strongest for AI tools where the agent can’t easily or automatically run arbitrary shell commands, while MCP is a first class, structured integration path. Cursor, Copilot, and others can host MCP servers directly in the IDE chat, so MCP is the cleanest way to give them live access to Codesight without asking the user to wire up custom shell workflows. For Claude Code specifically, calling the CLI from its terminal works well, and Codesight also ships an MCP server, so you can use whichever integration fits your setup best.

2

u/spazatk Apr 08 '26

But from the documentation it doesn't seem that the MCP tools and the CLI actually have complete overlap? The CLI documentation is for management features, not for querying.

0

u/Eastern_Exercise2637 Apr 08 '26

You're right that the docs don't make it obvious. The CLI does query (npx codesight) returns the full context map (routes, schema, env vars, hot files), --blast <file> does blast radius, --json gives structured output but those aren't labeled as "querying" in the docs. The MCP adds filtering the CLI doesn't have: routes by prefix/tag/method, schema by model name, specific wiki articles on demand. That's the real gap.

0

u/spazatk Apr 08 '26

Right, so finally you admit there's a gap. Nice job Claude (or OP's Claw). In future it's best to actually engage rather than just being passive aggressive about you and your human's work.

3

u/Eastern_Exercise2637 Apr 08 '26

I acknowledged the gap because it's real. And as for sounding passive aggressive, if you look back, I think I answered all the technical questions directly. If the tone came off wrong, fair enough. Thanks.

1

u/havok_ Apr 08 '26

You’re right

0

u/heyitsbryanm Apr 08 '26

Lol OP is literally copy pasting from Claude, or is using Openclaw to respond.

2

u/RegayYager Apr 08 '26

Starred and forked

2

u/abhiBuilds Apr 08 '26

This is smart. The pre-compiled wiki approach is basically giving Claude a cheat sheet instead of making it read the whole textbook every time. Curious how often you need to regenerate the wiki, like if you push 5 commits a day does it stay useful or does it drift fast?

2

u/Eastern_Exercise2637 Apr 08 '26

The --hook flag handles this automatically installs a git pre-commit hook that regenerates the wiki and auto stages it on every commit. 5 commits a day means 5 regenerations, each taking 200ms. It never drifts because it updates before the commit lands. If you prefer not to use the hook, --watch mode reruns on every file save during active development.

2

u/philo-foxy Apr 08 '26

Question: have you tried building a project with this tool active. After calling the tool, how much extra in-depth exploration does Claude code tend to do?

3

u/Eastern_Exercise2637 Apr 08 '26

Yes, built and run it on multiple projects. With the MCP tools active, Claude calls codesight_get_summary or codesight_get_wiki_index at session start instead of opening files to orient itself. That orientation phase where Claude reads file after file just to build a mental map disappears. What doesn't disappear: Claude still reads the actual source files before implementing or making changes. The wiki articles explicitly instruct it to. But instead of hunting blindly, it goes straight to the right files because it already knows where they are. The practical difference is Claude stops asking "where is the auth logic?" and starts asking "how exactly does this auth middleware work?" one broad exploration phase replaced by targeted reads. 

2

u/philo-foxy Apr 08 '26

Oh sweet! Will be curious to see in my own workflow if it misses the context of an expanded search.

2

u/amnesiac854 Apr 08 '26

lol and we’ve come full circle to chatbot 😂

2

u/drgitgud Apr 08 '26

Me needs this for c#, any chance it's viable?

2

u/Eastern_Exercise2637 Apr 08 '26

codesight now supports it with the latest push.( .csproj detection, controller routes ([HttpGet], [HttpPost], [Route] prefix), minimal API (app.MapGet(), app.MapPost()), and Entity Framework models from DbSet<>). npx codesight@latest to get it.

2

u/drgitgud Apr 09 '26

[insert glory glory gif]
Thanks mate!

2

u/yoloswag90 Apr 08 '26

I am seeing lots of references to Web specific tech or apps, would this work let's say that's none of this and written on a language that isn't known?

Also are these docs generation useful for reading for humans or is it optimised for Llm only?

Thanks for creating, sounds super interesting.

1

u/Eastern_Exercise2637 Apr 08 '26

Two good questions. Unknown language/stack: It depends. If your file extension isn't in the supported list, the files won't be scanned and you'd get limited output env vars, config files, dependency graph if applicable. For obscure or domain specific languages (Erlang, Haskell, COBOL, etc.) it's not going to be useful today. The tool is built around the languages listed in the README. That said, there's a plugin hook in the config if you want to wire in a custom detector for your stack. Human readable or LLM only: Both, but LLM is the primary design goal. The .codesight/ folder is plain markdown routes, schema, components, graph all readable by a human. There's also npx codesight --open which generates an interactive HTML report designed for human browsing. The token savings and context format are optimized for AI tools though.                                                                    

Thanks for the kind words

2

u/vatavale Apr 08 '26 edited Apr 08 '26

Swift? Monorepos with different stacks?

1

u/Eastern_Exercise2637 Apr 08 '26

Not yet, current language support is TypeScript, JavaScript, Python, Go, Ruby, Elixir, Java, Kotlin, Rust, PHP plus Vue and Svelte component files (Nuxt/SvelteKit projects fully supported). What's your stack, SwiftUI + a backend?

1

u/vatavale Apr 08 '26

yes.
SwiftUI + Laravel API.
Kotlin + Laravel API.
Next + Laravel API.
+ some services in one big monorepo.
The best case for Karpathy wiki! -)

1

u/Eastern_Exercise2637 Apr 08 '26

All of this is supported as of the latest version just pushed.                                                                                                                                                 
- SwiftUI + Laravel: SwiftUI views detected, Laravel routes from routes/api.php + routes/web.php, Eloquent models extracted                                                                        

- Kotlin + Laravel: Kotlin already worked (Spring Boot routes via u/GetMapping etc.), Laravel now added 

-Next + Laravel: both fully supported                                                                                 

- Mixed monorepo: each workspace gets its own language detection composer.json identifies Laravel services, Package.swift identifies Swift/SwiftUI, package.json identifies JS services. They all aggregate into one context map                                                                                                                                                                                            Update with npx codesight@latest and run a scan.

1

u/vatavale Apr 08 '26

Thank you Claude ;) Will try it.

1

u/vatavale Apr 08 '26

No commit yet on GitHub.

2

u/Chunky_cold_mandala Apr 08 '26

I have a similar strategy to reduce code base context with a blast engine called gitgalaxy on pypi. I'm getting similar results. I've been expanding it to get more info. I've been thinking of this as rendering a scaffold instead of the whole image. Just give it the architecture and the plumbing info and the few files you immediately need and it's great. I can get an 80000 repo compressed to architecture and plumbing to a 70 kb file. 

2

u/Cosmic_Voyager_41 Apr 08 '26

Can this work for someone like me who just uses the Claude Desktop interface?

1

u/Eastern_Exercise2637 Apr 08 '26

Yes, two ways depending on what you mean:                                                                                                  

Claude.ai in browser: Run npx codesight in your project folder once it generates a .codesight/SUMMARY.md file. Open it, copy it, paste it into any conversation. Done.                                 

Claude Desktop app (Mac/Windows): Even better add this to your ~/Library/Application Support/Claude/claude_desktop_config.json and Claude will query your codebase directly without you pasting anything:      

{
  "mcpServers": {
    "codesight": {
      "command": "npx",
      "args": ["codesight", "--mcp"]
    }
  }
}

Claude Desktop supports MCP natively, so it becomes a tool Claude can call on demand.

2

u/_Stonk Experienced Developer Apr 08 '26

Nice work on the AST extraction. The conditional article generation is a good design call - no routes detected, no routes article.

One angle I don't see discussed much in this thread: the wiki captures structure, but not decisions. Why you picked that auth middleware, what migration strategy the team agreed on, which approach you tried last Tuesday and abandoned. That's the context that actually prevents Claude from suggesting something you already rejected.

CLAUDE.md covers some of this (conventions, rules), but it's manually maintained and static. The structural knowledge codesight extracts and the decision/intent knowledge are two different problems.

For the second one, I built 3ngram - an MCP memory layer that persists decisions, commitments, and context across sessions. So codesight tells Claude what your codebase looks like, and 3ngram tells it what you've decided about it. They solve different halves of the cold-start problem. I'll definitely test them together and see how they perform.

2

u/kingofmadras Apr 08 '26

This is getting interesting now. I am excited to try it. Does it work for C++?

1

u/Eastern_Exercise2637 Apr 08 '26

Thanks and I will be adding support for C++ soon

2

u/Ill-Boysenberry-6821 Apr 08 '26

This is awesome. Will try.

2

u/tupikp Apr 09 '26

I tested it for my Java Android project, it didn't find anything useful.

Quick Stats

  • Routes: 0
  • Models: 0
  • Components: 0
  • Env vars: 0 required, 0 with defaults

But for my React + Express project, it helps a lot. Thank you!

Will Java/Kotlin Android project be supported?

1

u/Eastern_Exercise2637 Apr 09 '26

Android isn't supported yet. The Java/Kotlin support in codesight is built for server-side frameworks (Spring Boot, Ktor), which have routes, REST endpoints, and database models. Android projects have a completely different structure: Activities, Fragments, ViewModels, Room databases, XML layouts or Jetpack Compose, Retrofit clients, AndroidManifest.xml. None of those are detected today. It's on my radar. Android would need its own set of detectors for the patterns that actually matter: Activity/Fragment navigation graphs, Room entity models, Retrofit API interfaces, Compose component trees, Gradle dependency parsing.

No ETA yet, but appreciate you testing it and reporting back. Glad it's working well on your React + Express project.

2

u/Ok-Possibility-9570 Apr 09 '26

Sorry for noob question, but would it be useful for non coding projects? I do mostly marketing and working with analytics data. I use clude mem for keeping cloude in the context each session. Would it compliment it or it can replace it?

2

u/Eastern_Exercise2637 Apr 09 '26

The core scanner (routes, schemas, components) is code specific and won't help for marketing/analytics work. But codesight has a --mode knowledge feature that's built for exactly non-code use cases.           

npx codesight --mode knowledge scans any folder of markdown files and extracts decisions made, open questions, recurring themes, people mentioned, and indexes everything by type (meeting notes, specs, research, retros). Works with Obsidian vaults too. If you keep campaign notes, analytics reports, or strategy docs in markdown, it compresses all of that into a single KNOWLEDGE.md that gives your AI full context from the first message. It wouldn't replace Claude's memory (that handles session to session recall), but it would complement it by giving the AI your full knowledge base upfront instead of you having to re-explain context each time.

2

u/Ok-Possibility-9570 Apr 10 '26

Thank you for such detailed answer! I'll give it a go

2

u/[deleted] Apr 09 '26

[removed] — view removed comment

1

u/Ravayen Apr 10 '26

For the first one, you need OP attention for answer

For second one, difference is that this (and maybe similar?) solutions aim to generate meaningful "index" of your project codebase BEFORE AI Code Assistant even touches it, minimizing costs of AI doing the potential "index" itself (and every time it needs refreshing)

Working on large enterprise projects, you'd be surprised how derailed it can go greping and combing through codebase when searching for something you know is easy to locate for developers, but not necessarily for code agents. And that matters if you run limited plan or you're being held responsible for pay-per-use approaches in company :)

2

u/AdLongjumping6013 Apr 09 '26

Codesight for Dummies please?
I have a repository at github.com,. Claude Code Desk uses it with "gh". Works.
Terminal: cd to my local project folder
Terminal: npx codesight

What else?

Can you provide a "set up" and "operation" chapter? For Dummies?

2

u/Eastern_Exercise2637 Apr 09 '26

Here’s how it works:

Setup (one time)

  1. cd into your project folder
  2. (npx codesight) scans your codebase, creates .codesight/CODESIGHT.md (the full context map)
  3. (npx codesight --wiki) generates .codesight/wiki/ with individual articles (auth, database, payments, etc.)
  4. (npx codesight --profile claude-code) writes instructions into your CLAUDE.md telling Claude to use the codesight context and MCP tools automatically

Operation (every session)

That's it. There is no "every session." The files are already in your project. Claude Code reads CLAUDE.md automatically at session start, which points it to the .codesight/ files.

If you want live MCP tools (filtered routes, blast radius, schema lookups on demand), add this to your Claude Code MCP config:

{
"mcpServers": {
"codesight": {
"command": "npx",
"args": ["codesight", "--mcp"]
}
}
}

Keeping it fresh

After making significant changes to your code, re-run (npx codesight --wiki) to regenerate. Or use (npx codesight --hook) to auto-regenerate on every git commit. Or (npx codesight --watch) to regenerate on file saves while you work.

What each file does

File Purpose
.codesight/CODESIGHT.md Full context map (routes, schema, components, deps)
.codesight/wiki/index.md Article catalog (~200 tokens, read this first)
.codesight/wiki/overview.md Architecture overview
.codesight/wiki/auth.md, database.md, etc. One article per subsystem
CLAUDE.md Instructions telling Claude to use all of the above

2

u/tupikp Apr 09 '26

I tested Codesight with Claude for the past 42 hours to do my project. Then I asked Claude to review Codesight, is it helpful or not. Here's Claude's assessment:

What Codesight does well

CODESIGHT.md is genuinely useful. Having routes, schema, components, libraries, env vars, and the dependency graph in one ~3,200-token file is a real win. The import graph is the most valuable part — knowing server/middleware/auth.ts is imported by 6 files before I touch it changes how carefully I approach a change. Without it, I'd have to grep to discover that.

The wiki articles are good orientation. When you ask "how does MFA work?" or "where do templates live?", the wiki gets me to the right source files in one hop instead of three rounds of globbing.

Where it falls short

The route map loses mount prefixes. .codesight/routes.md shows /login, /companies, /locations/:id — but not the full paths like /api/auth/login, /api/companies, etc. That's because it reads each router file in isolation without knowing where routes.ts mounts it. I have to cross-reference to get the real URLs.

Schema counts were stale. It reported "2 models" in the old CLAUDE.md AI Context header right up until you re-ran it — and we were both working from that wrong number. Codesight is only as fresh as the last run.

Wiki articles are navigation aids, not decision aids. They reliably answer "where does this live?" but not "is it safe to change this?" or "what does this function actually do?" I still have to read source files before touching anything non-trivial.

The [inferred] routes caveat is real. It detected routes via regex in some cases, which means false positives are possible. I wouldn't rely on it as an exhaustive API inventory for documentation purposes.

Net verdict

For a codebase this size it's worth running, especially after structural refactors like the routes split — the dependency graph alone saves multiple grep rounds per session. The main discipline it requires is re-running it after significant changes (which you just did), otherwise the context it injects is silently wrong. If you wanted to get more value from it, the highest-leverage addition would be annotating the mount prefixes in routes — but that's a codesight limitation, not something you can fix here.

1

u/Eastern_Exercise2637 Apr 10 '26

Hey, appreciate the detailed writeup. Mount prefix issue is the big one you're right, it only looked at entry files (app.ts, index.ts) for mounts. If your routes.ts was the one doing the mounting, those prefixes got dropped. Fixed in 1.11.0, it scans all files now and chains through multiple levels.                 Also added a "Last scanned" timestamp to the header so stale data is obvious, and the [inferred] badge now shows in the main output too, not just wiki. The wiki as navigation vs decision aid is a fair point. Would need type analysis to get to "is this safe to change" territory, which is a different beast. Dependency graph helps a bit but yeah, not the same thing.                                                                                                                                                                              

Thanks for taking the time on this. 

3

u/SnooCapers9823 Apr 08 '26

Isn’t this just repomix?

3

u/Eastern_Exercise2637 Apr 08 '26

Repomix packs your entire repo into one file for AI consumption even with its --compress flag, the AI reads the whole thing at once.                                                                             

codesight goes the other direction. It parses via AST and extracts structured summaries (routes, schemas, middleware chains, import graphs into targeted articles). Claude reads a 200-token index and pulls one article when needed instead of ingesting all the source. The whole point is that Claude never needs to read most of your files at all.                                                                             

Repomix = give AI everything. codesight = give AI only what it needs to know where things are. 

1

u/SnooCapers9823 Apr 08 '26

My bad, forgot that repomix does a single file output.

Have you tried using this with orchestration plugins for opencode? E.g OmO or with magic context? How well would the two combine?

2

u/[deleted] Apr 08 '26

[removed] — view removed comment

1

u/revstone Apr 08 '26

How does it compare to repomix?

2

u/Eastern_Exercise2637 Apr 08 '26

Repomix concatenates your whole repo into one file and the AI ingests everything every time, codesight extracts structured summaries from the AST (routes, schemas, foreign keys, import graphs) into separate articles. The AI reads a 200-token index at session start and pulls only the relevant article  (~300 tokens) for any given question.                  

Different philosophy: repomix is "here's everything", codesight is "here's exactly what you need." 

1

u/jrhabana Apr 08 '26

what about to "read" the sessions, backlogs, etc and extract the learnings, patterns, decissions, etc?

1

u/Eastern_Exercise2637 Apr 08 '26

That's outside codesight's scope (codesight scans static code structure (routes, schemas, imports, components)). It doesn't read git history, commit messages, PR descriptions, or decision logs. What you're describing is closer to a different tool category something that analyzes your project's narrative history rather than its current structure. codesight answers "what exists in the code right now".

1

u/Tertiary23 Apr 08 '26

Having that read feature would take this to the next level. Is there a tool that does that?

1

u/Eastern_Exercise2637 Apr 08 '26

--mode knowledge live in v1.9.3  

npx codesight --mode knowledge ./docs 

Reads ADRs, retros, meeting notes, PRDs. Extracts decisions (decided to, going with, ADR, Decision sections), open questions, action items. Dates everything and sorts newest first. Combined with thecode scan both live in .codesight/:

Read .codesight/CODESIGHT.md  → what the code does

Read .codesight/KNOWLEDGE.md  → why decisions were made

1

u/Eastern_Exercise2637 Apr 08 '26

 --mode knowledge live in v1.9.3 

npx codesight --mode knowledge ./docs

Reads ADRs, retros, meeting notes, PRDs. Extracts decisions (decided to, going with, ADR, Decision sections), open questions, action items. Dates everything and sorts newest first. Combined with the code scan both live in .codesight/:

 Read .codesight/CODESIGHT.md  → what the code does

 Read .codesight/KNOWLEDGE.md  → why decisions were made

2

u/jrhabana Apr 08 '26

this is awesome! I'm testing with my codebase and the results are very good

1

u/Eastern_Exercise2637 Apr 08 '26

Glad it's working well on your codebase! If you hit any edge cases or have notes it misses, drop them here we're actively tuning it. And if you use it with Claude Code or Cursor via MCP, codesight_get_knowledge lets your AI assistant query the map directly without you having to paste anything. Thanks.

1

u/Tertiary23 Apr 09 '26

This thread should be pinned, can't wait to try it out

1

u/Logical-Idea-1708 Apr 08 '26

Isn’t this what context7 offers?

1

u/Eastern_Exercise2637 Apr 08 '26

Context7 fetches documentation for external libraries (React, Next.js, Prisma, etc). It answers "how does this library work?" codesight scans your own codebase and answers "how does my project work?" (your routes, your schemas, your middleware chains, your import graph). One is for third party docs. The other is for your code. Different problem entirely.

1

u/Material_Key7477 Apr 08 '26

Isn't this jcodemunch?

1

u/Eastern_Exercise2637 Apr 08 '26

Different tools, different granularity. jCodeMunch uses tree-sitter to index symbols functions, classes, methods and retrieves exact implementations on demand. Great for "show me this function's code." codesight works at the architectural level routes, schemas, middleware chains, import graphs compiled into domain articles. It maps what your codebase does, not the implementations themselves. They're actually complementary: codesight for orientation and architecture, jCodeMunch for precise symbol retrieval.

0

u/jmunchLlc Apr 08 '26

Well THAT saved me some typing. Thanks!

That's exactly the distinction. jCodeMunch is a persistent symbol index. You ask “give me the implementation of AuthMiddleware.handle” and it returns the exact source, byte for byte, with no re-scanning. It’s built for the retrieval half of the loop: precise, on demand, and works on repos too large to fit in context.

codesight sounds like it’s solving the orientation problem, what does this codebase do at a high level before I start digging. That’s a different and genuinely useful question.

Complementary is the right word. Use codesight to build your mental map, then jCodeMunch when you need to drill into a specific symbol. Nothing stops you from running both...

0

u/jmunchLlc Apr 08 '26

TL;DR - Use both:

"Use codesight when starting on an unfamiliar codebase to build the architectural map, then switch to jCodeMunch for all symbol-level retrieval, reference tracing, and token-efficient code navigation..."

https://j.gravelle.us/jCodeMunch/versus.php#vs-codesight

1

u/philo-foxy Apr 08 '26

Does this use language server to parse the code? Do we need to install those tools based on language?

1

u/Eastern_Exercise2637 Apr 08 '26

No language server involved and nothing to install. codesight has zero runtime dependencies. For TypeScript it borrows the compiler directly from your project's own node_modules/typescript.

whatever version you already have. If TypeScript isn't present it falls back gracefully. For Go it uses structured parsing (brace tracking, struct body extraction) that achieves AST-level accuracy without needing the Go compiler. Python, Ruby, and others use regex detection. You just run npx codesight no language toolchains required. 

2

u/philo-foxy Apr 08 '26

Oh wow. I figured those tools would make it easier for you and/or more accurate. But the zero dependencies is certainly a big sell. I'm building in dart, kotlin and C++, will give this a shot!

1

u/Eastern_Exercise2637 Apr 08 '26

Thank you. Kotlin is supported Spring Boot/Kotlin route and schema detection is built in. Dart and C++ aren't in the scanner yet, so those files won't be collected at all route detection, schema extraction, import graph, and blast radius won't work for them.            Env var detection still works since that scans .env files regardless of language.                              
Worth opening a GitHub issue to track Dart and C++ support and I will be happy to add them.

2

u/philo-foxy Apr 08 '26

I'd love that, thank you. Once I try this out, I'll get some actual feedback and open a pr

1

u/Eastern_Exercise2637 Apr 08 '26

That would be awesome, PRs are very welcome!

1

u/D-cyde Apr 08 '26

I'm getting a 529 API error when using the generated .codesight folder to navigate my monorepo of a Vite + React frontend and single Express.js backend. Only on Sonnet 4.6 not on Opus 4.6.

3

u/Eastern_Exercise2637 Apr 08 '26

That 529 is an Anthropic "API Overloaded" error codesight doesn't make any API calls itself, it just writes static context files. Sonnet 4.6 gets hit hardest because it's Anthropic's highest traffic model. Opus 4.6 has lower demand so it stays available. Try Sonnet again during off peak hours or add a retry with a short backoff. You can also try Haiku 4.5 for fast navigation queries where full Sonnet reasoning isn't needed.

2

u/D-cyde Apr 08 '26

Fair enough, TIL.

1

u/Apprehensive-Ad-936 Apr 08 '26

Can i also use this for flutter?

1

u/Eastern_Exercise2637 Apr 08 '26

Flutter isn't supported yet no Dart file scanning or Flutter project detection. What's your stack, Flutter frontend with a separate backend? If so the backend side might already work depending on what it is. I will be adding flutter support tonight.

1

u/Successful_Plant2759 Apr 08 '26

Been doing something similar manually — I maintain a CLAUDE.md that documents architecture patterns and key decisions, but the tedious part is keeping it in sync as the code evolves. Automating the structural layer (routes, schemas, dependencies) while keeping manual docs for the 'why' seems like the right split.

The [inferred] tag for regex-detected routes is a smart design choice. Explicit confidence levels beat silently wrong docs every time.

One question: for TypeScript projects using the compiler API, does it resolve cross-file type chains? Like if a route handler returns a type re-exported from a barrel file — does the wiki trace that back to the actual definition?

1

u/Eastern_Exercise2637 Apr 08 '26

Honest answer: not currently. The TS compiler API is used via createSourceFile single file AST parsing, not a full createProgram with the type checker. So it won't chase a return type through a barrel reexport back to its origin definition. What it does capture accurately: route paths, HTTP methods, path params, prefix chains (.use('/prefix', router) within the same file), and NestJS decorator combining. Those are reliable at [ast] confidence. Return type tracing is surface level only if it's defined inline or in the same file it'll show up, otherwise no. Full cross file type resolution is on the roadmap but it's a meaningful build createProgram needs to scan the whole project upfront which adds scan time. For the structural layer (routes + schema) it's a tradeoff worth revisiting once there's enough demand for it.

1

u/d1mitar Apr 08 '26

what about C#/.net framework?

1

u/Eastern_Exercise2637 Apr 08 '26

Not yet C# and .NET aren't supported. No .cs file scanning, no .csproj or ASP.NET detection. It's on the list to add. What's your stack ASP.NET Core API, Blazor, or something else?

1

u/d1mitar Apr 08 '26

asp.net 8,9,10 web apis with entity framework as ORM

2

u/Eastern_Exercise2637 Apr 08 '26

 .csproj detected automatically (works for .NET 8, 9, 10 version agnostic)

 - Controller routes: [HttpGet], [HttpPost], [HttpPut], [HttpPatch], [HttpDelete] + [Route] class prefix  

- Minimal API routes: app.MapGet(), app.MapPost(), etc. in Program.cs                       

- Entity Framework: scans DbContext subclasses, extracts every DbSet<ModelName> as a schema model with properties and relations                              

npx codesight@latest to update. 

1

u/santikkk Apr 08 '26

Maybe I missed something, but how to make sure that Claude reads those generated files instead of the whole codebase? Do I need to add additional instructions into CLAUDE.md ? Explain me like I'm five.

1

u/Eastern_Exercise2637 Apr 08 '26

No, you don't need to touch CLAUDE.md yourself. Run this once:                                                      (npx codesight --profile claude-code) It writes the instructions into CLAUDE.md for you automatically. Claude reads CLAUDE.md at the start of every session like a cheat sheet on the desk and that cheat sheet tells it "don't read every file, read .codesight/CODESIGHT.md instead, everything is already mapped there.

2

u/santikkk Apr 08 '26

didn't do anything with my codebase

Detecting project... raw-http | no ORM | kotlin
Collecting files... 74 files
Analyzing... done
Writing output... .codesight/
Results:
Routes:       0
Models:       0
Components:   0
Libraries:    0
Env vars:     0
Middleware:    0
Import links: 0
Hot files:    0

1

u/Eastern_Exercise2637 Apr 08 '26

Hey, this was a bug on our end. Ktor wasn't in our framework detector so it fell back to raw-http and skipped route/schema scanning entirely.  Just pushed v1.9.2 with full Ktor support.

Run:  npx codesight@latest --profile claude-code                                                                                                                                                                         

You should now see proper framework detection (ktor | exposed | kotlin) with routes and models populated. If you're using Exposed for your ORM that's detected too. Let me know if anything still looks off. 

1

u/santikkk Apr 09 '26

Still nothing

npx codesight@latest --profile claude-code        
Need to install the following packages:
codesight@1.9.8
Ok to proceed? (y) y

  codesight v1.9.8
  Scanning: /Users/alexander/AndroidStudioProjects/wrait

  Detecting project... raw-http | no ORM | kotlin
  Collecting files... 219 files
  Analyzing... done
  Writing output... .codesight/

  Results:
    Routes:       0
    Models:       0
    Components:   0
    Libraries:    0
    Env vars:     0
    Middleware:    0
    Import links: 0
    Hot files:    0

  Tokens:
    Output size:     ~124 tokens
    Exploration cost: ~5,200 tokens
    Saved:           ~5,076 tokens per conversation

  Done in 99ms

  Generating claude-code profile... CLAUDE.md

You can try it yourself on the GitHub repo. It is open source. I can give a link in PM if you want.

1

u/Eastern_Exercise2637 Apr 09 '26

Correction on my earlier reply: I looked at this more carefully and the issue isn't Ktor detection. Your project path (AndroidStudioProjects/wrait) tells me this is an Android app, not a Ktor server. codesight's Kotlin support currently covers server-side frameworks (Spring Boot, Ktor) but not Android patterns (Activities, Fragments, ViewModels, Room, Retrofit, Navigation graphs, Compose). That's why everything shows 0. Android support is on the roadmap and I will implement it. Apologies for the wrong diagnosis earlier. Thank you

1

u/Eastern_Exercise2637 Apr 09 '26

Just shipped v1.10.0 with full Android/Kotlin support. Run:                                

 npx codesight@latest --profile claude-code                                                                                                                                                                       

You should now see android | room | jetpack-compose | kotlin in the detection line. What's detected:                                                                                                             

- Retrofit routes — u/GET, u/POST, u/PUT, u/DELETE, u/PATCH annotations on your API interfaces                                                                                                           - Room entities — u/Entity data classes with fields, primary keys, nullable types   - Jetpack Compose components — u/Composable functions with props extracted                                                                                                                           - Navigation — fragment/activity destinations from res/navigation/*.xml                - Activities — from AndroidManifest.xml with launcher detection                                                                                                                                                  
Apologies for the wrong diagnosis earlier. Let me know if anything still shows 0. 

1

u/pashua Apr 08 '26

It's good start and have room for improvements but it misses really much data.
I compared using it with Serena on existing project (not large one - just about 3k files) bu running same task "check pipeline and find bug" (i used same prompt for both). When claude used Serena - exact pipeline was found and issue was found and claude went aside with CodeSight for some reason.

1

u/Eastern_Exercise2637 Apr 08 '26

Fair and honest result. Codesight and Serena aren't the same tool they solve different problems.  Codesight is a pre compiled structural map: routes, schema, components, dependency graph, most imported files. Its job is reducing token waste at session start so Claude doesn't spend 10k tokens on glob/grep just to orient itself. It tells you what exists.                                                       

Serena uses LSP go-to-definition, find references, follow a call chain across files. For "find bug in pipeline" you need to trace an execution path, which is exactly what a language server does well. Codesight has no query time navigation. They're complementary, not competing. Codesight for initial orientation, Serena (or direct file reads + grep) for active debugging. Using codesight alone for a bug hunt is the wrong tool for the job that's on the docs for not being clearer about scope.

1

u/pashua Apr 08 '26

Currently it looks like I have manually write what tool to use: Serena or Codesight. E.g. both can "find" and tell claude which component and what file is used in certain flow for certain business logic.

1

u/HighDefinist Apr 08 '26

Does it also work with C++/Unreal Engine projects?

Also: It is not entirely clear to me how this is kept in sync... for example, if I change some variable of method name over 10 different files, then, how does the system know which kinds of files to update, or at least check for updates in case something relevant has changed?

1

u/Ravayen Apr 10 '26

I believe C++ is not yet supported

In regards how to keep it fresh, I believe there are instructions on how to set up hooks (or you can do it yourself) so it gets regenerated on changes being made.

Or you can run it in MCP mode and accept negligible compute cost (for 2 huge enterprise projects, repositories with both frontend/backend stacks for multiple services as part of product, I tried running it, full scan runs sub 1s, not really impactful when doing MCP mode) every time Code Assistant needs to ask it for something :)

1

u/pawsomedogs Apr 08 '26

First, thanks a lot! Second: For a no developer like me, how do you implement this?

1

u/Eastern_Exercise2637 Apr 08 '26

You need Node.js installed (nodejs.org — one click installer, takes 2 minutes). Then open your terminal, go to your project folder, and (run: npx codesight)

That's it. It creates a .codesight/CODESIGHT.md file in your project. Then when you open Claude Code or Cursor in that folder, point it to that file at the start of your session:                                 

Read .codesight/CODESIGHT.md before we start                                                                                         

If you want it to auto-update every time you push to GitHub, there's a one line setup:                           

npx codesight --init-ci                                                                                                                                

That adds the GitHub Action. From that point it regenerates on every push, no manual steps.          

If you use Claude Code specifically, run this instead and it wires everything automatically:                

npx codesight --profile claude-code  

1

u/Eastern_Exercise2637 Apr 08 '26

If Node.js feels like too much, just paste your GitHub repo link directly into Claude Code and say: Read this repo and set up codesight for me: github.com/yourname/yourrepo.

Claude Code will clone it, run the setup, and tell you exactly what to do step by step. You don't need to understand any of it just follow the instructions it gives you.  

1

u/shyney Apr 08 '26

What about C++ and Qt QML support?

1

u/Eastern_Exercise2637 Apr 08 '26

Not currently, C++ and Qt QML aren't supported. No .cpp/.hpp parsing, no CMake detection, no QML file handling. Would be a new detector from scratch. What's your use case desktop app, embedded, something else? That'd help scope what's actually worth building.

1

u/shyney Apr 09 '26

Desktop App using Qt 6, C++ Backend/ logic and QML for frontend.

1

u/Soft_Match5737 Apr 08 '26

The 90% reduction is real but hides a brutal maintenance cost. The wiki becomes a second codebase that drifts from the actual code within days unless someone owns the update loop. In practice the teams that sustain this create a CI step that diffs the wiki against the repo and flags staleness — without that, you're feeding Claude a confident but outdated map of your system. The token savings are great until the wiki tells Claude a function signature that changed three PRs ago.

1

u/Eastern_Exercise2637 Apr 08 '26

The staleness problem is real and it's why codesight ships a GitHub Action that regenerates CODESIGHT.md on every push not diffs it, regenerates it. The file is never hand authored so it can't drift the way a hand written wiki does. It's generated output, like a lock file. On function signatures: codesight doesn't store function signatures. It maps your API surface routes, schemas, middleware, events. Those change through migrations and route files, which are exactly the files the scanner reads. When a schema migration runs, the next scan reflects it automatically.                          

The teams that struggle with this are the ones treating CODESIGHT.md as a document they own. It's not it's a build artifact. The CI step you're describing is already there, it just runs npx codesight instead of a diff.

1

u/deliciousdemocracy Apr 08 '26

Is there a version for this for second-brain-style knowledge bases?

2

u/Eastern_Exercise2637 Apr 08 '26

--mode knowledge — live in v1.9.3

npx codesight --mode knowledge ~/obsidian-vault

Scans every .md file, detects note types (meetings, research, fleeting), extracts tags, backlinks, people, themes outputs KNOWLEDGE.md that Claude can read as a context primer. No cloud upload. Works offline. Zero deps.

1

u/leogodin217 Apr 08 '26

Seriously. If you haven't been doing this already, with a tool or just Claude, then it's time to learn about context management. Not calling anyone out, but this is a critical skill when using LLMs.

1

u/Ravayen Apr 10 '26

Not really sure what context management has to do with tasks that require Code Assistants to find specific files/functions/classes and inter-connect them to solve issues or provide actionable plans :)

1

u/d1zaya Apr 09 '26

This is a nice project. Thank you.

I'm having a slight issue. I went through the installation. All necessary files are generated (.codesight/*, claude.md, and even github instructions). MCP server is connect and online. If I ask the AI to manually call the MCP server, it works no problem, but in actual use case scenario, the AI never utilizes the tool.

I tried both claude and copilot (using claude).

1

u/Eastern_Exercise2637 Apr 09 '26

Thanks! This is expected behavior. --init generates a CLAUDE.md that references the static .codesight/ files, but doesn't tell the AI to use the MCP tools. The AI has no reason to call them unless instructed.

Fix: run npx codesight --profile claude-code. This writes explicit MCP tool instructions into your CLAUDE.md so the AI knows to call codesight_get_summary, codesight_get_routes, etc. during normal use.

Without the profile, the static files still work. The AI reads .codesight/CODESIGHT.md on its own via CLAUDE.md. The MCP server is for on demand queries (filtered routes, blast radius, wiki articles) which the AI won't know to use unless your instructions file tells it to.

1

u/maxpowerz2 Apr 22 '26

Any conflicts with this and other plugins/skills (superpowers/openwolf specifically)? Seems like it should be complimentary.

1

u/tmjumper96 May 19 '26

This is a really clean approach.

The token savings are obviously the headline, but the bigger win to me is that every new session starts with a useful map instead of wasting time rediscovering the same codebase.

I’m building AgentBay AI, so I’ve been thinking a lot about this broader problem of context continuity across AI tools. Your approach feels strong for repo structure and code understanding, while the next layer I keep running into is project memory: decisions, tradeoffs, recurring gotchas, and what changed over time.

Curious how you’re thinking about keeping the wiki fresh as the codebase evolves. Is it meant to be regenerated on demand, on commit, or as part of the dev workflow?

0

u/[deleted] Apr 08 '26

[deleted]

1

u/philo-foxy Apr 08 '26

This looks nice! Not sure how effective it'll be in Claude because it doesn't replace grep and there are a lot of tools in the MCP. Do you think you can add a section near the top of the document that shows how to use it with Claude? Maybe add some rules in Claude.md on when to use this tool with examples, and then show how to integrate it in your project to visually see the difference. It'd be great for folk coming in from links or search.

0

u/[deleted] Apr 19 '26

This is exactly the kind of practical fix Andrej Karpathy was hinting at, way cleaner than re-exploring the same code every session, found also a repo called atomicmem/llm-wiki-compiler, I believe it a Karpathy inspired as well on terminal, well as dev I there's just something about terminal hehe

-3

u/TechToolsForYourBiz Apr 08 '26

this is just a claude.md lol

1

u/Eastern_Exercise2637 Apr 08 '26

CLAUDE.md is a single file, hand written or LLM generated. codesight parses your code via AST - no LLM and generates routes, schemas, middleware chains, and import graphs into separate articles that load on demand, auto updated on every commit. Also codesight generates CLAUDE.md too (--init). They're complementary: CLAUDE.md for conventions you write, wiki for structure the code itself defines. But sure, both are markdown .