[READ-ONLY] Mirror of https://github.com/agbocsardi/academic-lsp. Experimental language server diagnostics for academic prose
README.md

academic-lsp #

Experimental language server diagnostics for academic prose.

Lore #

Inspired by Martin Kleppmann's Bluesky post asking for IDE-style red squiggles when prose refers to a non-obvious concept before defining it:

https://bsky.app/profile/martin.kleppmann.com/post/3mn5fqwajvs2l

One of the professors at Gergő's school says that "academic writing is more code than literature" because it has to be structured immensely. This project takes that seriously: if prose has code-like structure, it should be possible to build code-like diagnostics for it.

Goal #

Bring code-editor feedback loops to academic writing:

  • warn when an abbreviation is used before it is defined
  • warn when a potentially technical concept appears before a local definition
  • eventually warn when adjacent paragraphs do not connect cleanly

The first scaffold is intentionally small and deterministic. LLM-backed semantic diagnostics can come later once the editor/server loop is solid.

Current prototype #

The server currently emits diagnostics for common abbreviation patterns:

  • LSP used without a nearby earlier definition like Language Server Protocol (LSP)
  • abbreviation definitions are collected from the current document
  • diagnostics update through normal LSP textDocument/didOpen and textDocument/didChange events

Install for local development #

uv sync
uv run academic-lsp --help

Configuration #

The server is meant to be configured per project. Copy academic-lsp.example.toml to academic-lsp.toml and adjust it.

[rules]
files = [
  ".academic-lsp/rules/base.md",
  ".academic-lsp/rules/discipline.md",
  ".academic-lsp/rules/supervisor.md",
]

[llm]
provider = "litellm"
base_url = "https://opencode.ai/zen/go/v1"
model = "openai/deepseek-v4-flash"
temperature = 0.0
api_key_env = "OPENCODE_API_KEY"

API keys should live in environment variables, not in the project config.

Diagnostic triggers are configurable. Defaults are intentionally conservative:

  • deterministic checks run on idle by default
  • LLM-backed checks run on save by default
  • expensive checks should also support manual runs later
[diagnostics.abbreviations]
enabled = true
engine = "deterministic"
run_on = "idle"
debounce_ms = 1500

[diagnostics.paragraph_transitions]
enabled = true
engine = "llm"
run_on = "save"

This keeps typing responsive while still letting heavier semantic checks run when the draft reaches a stable point. The current LSP wiring runs deterministic diagnostics on open/change and LLM diagnostics on save when a model and API key are configured.

Rule files #

Rule files are plain Markdown so different disciplines, journals, supervisors, or writing courses can bring their own standards.

The repo includes two example rule packs:

  • rules/basic-academic-prose.md — small baseline rules used by tests and future evals
  • rules/jvgemert-writing.md — example adapted from Jan van Gemert's writing guidelines

Diagnostics should cite the rule that triggered them where possible, so warnings feel like configured writing standards rather than generic AI opinions.

Try it in Neovim #

This repo includes examples/smoke.md and a local academic-lsp.toml for testing.

export OPENCODE_API_KEY=...
nvim examples/smoke.md

Deterministic abbreviation diagnostics should appear after opening/editing. LLM diagnostics run on save.

Neovim sketch #

This can be registered as a custom LSP server. Diagnostics are normal LSP diagnostics, so Neovim can render them as inline virtual text — the ghost-text-style warning beside the prose.

vim.diagnostic.config({
  virtual_text = true,
  underline = true,
  signs = true,
  float = true,
})

vim.api.nvim_create_autocmd({ "BufReadPost", "BufNewFile" }, {
  pattern = { "*.md", "*.qmd", "*.tex" },
  callback = function()
    vim.lsp.start({
      name = "academic-lsp",
      cmd = { "uv", "run", "academic-lsp" },
      root_dir = vim.fs.root(0, { ".git" }) or vim.fn.getcwd(),
    })
  end,
})

If inline messages get too noisy, use virtual_text = false and rely on underlines plus hover/floats instead.

Direction #

Possible diagnostic layers:

  1. deterministic abbreviation checks
  2. noun-phrase/concept extraction with local definition tracking
  3. paragraph transition/coherence checks
  4. optional small local model pass for fuzzy academic-prose warnings

The design target is not autocomplete or rewriting. It is quiet, local, editor-native feedback while drafting.