<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://jon.recoil.org/atom.xml</id>
  <title type="text">Jon Ludlam&apos;s blog</title>
  <generator uri="https://github.com/xhtmlboi/yocaml" version="2">YOCaml</generator>
  <updated>2026-09-08T00:00:00Z</updated>
  <author>
    <name>Jon Ludlam</name>
  </author>
  <link href="https://jon.recoil.org/atom.xml" rel="self"/>
  <link href="https://jon.recoil.org" rel="alternate"/>
  <entry>
    <id>https://jon.recoil.org/blog/2026/09/summernotes.html</id>
    <title type="text">Summer notes</title>
    <updated>2026-09-08T00:00:00Z</updated>
    <published>2026-09-08T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Getting back to writing notes! Here&apos;s a dump of what&apos;s been going on. More detailed write-ups to follow at some point!&lt;/p&gt;
&lt;h2&gt;First class docs&lt;/h2&gt;
&lt;p&gt;The notion of making docs a &apos;first class&apos; feature of OCaml seemed pretty important to me, so I applied to present it at the OCaml Workshop, and it got accepted! I was really looking forward to going and discussing with everyone important in OCaml development how to progress this, and then life intervened and I found out that my daughter Skylark had a rather major surgery scheduled for exactly the day of the workshop. I stayed here to be with her, and happily it all went well. Meanwhile in France they played a recording I had made of my talk, and &lt;a href=&quot;https://choum.net/panglesd/&quot;&gt;Paul Elliot&lt;/a&gt; answered questions. The slides are &lt;a href=&quot;/static/talks/first_class_docs/firstclassdocs.html#5&quot;&gt;here&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;It has managed to start &lt;a href=&quot;https://github.com/ocaml/opam/discussions/7112&quot;&gt;a discussion&lt;/a&gt;, so we&apos;ll see where that ends up.&lt;/p&gt;
&lt;h3&gt;Odoc fixes&lt;/h3&gt;
&lt;p&gt;As part of this, I needed to make sure that installing the odoc files wasn&apos;t going to be just a massive waste of space. An initial investigation showed the odoc files were roughly as large as the entire lib directory - so installing them doubled its size. This was clearly unacceptable, so I spent some time looking at saving space. We&apos;d not even looked at this before, so I was pretty confident we&apos;d find some good space savings.&lt;/p&gt;
&lt;p&gt;There were 3 straightforward changes that made a big difference. Firstly, and most obviously, enabling compression. This was &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1486&quot;&gt;straightforward&lt;/a&gt; as OCaml itself has been compressing files since 5.1.0. Secondly, removing some optimisation in the representation of Identifiers that&apos;s become redundant since it was put in. This was a &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1479&quot;&gt;larger PR&lt;/a&gt;, but mostly mechanical. The one moderately interesting part of it was modifying the hash function of the identifiers to ensure we always traverse right to the &lt;code&gt;Root&lt;/code&gt; constructor in deeply nested identifiers. The last fix was just in the AST representing comments. In it, a &lt;a href=&quot;/reference/ocaml/odoc/odoc.model/Odoc_model/Comment/index.html#type-paragraph&quot;&gt;paragraph&lt;/a&gt; is an &lt;code&gt;inline_element with_location list&lt;/code&gt; - a list of inline_elements that each have a location - a start and end point. The &lt;a href=&quot;/reference/ocaml/odoc/odoc.model/Odoc_model/Comment/index.html#type-leaf_inline_element&quot;&gt;inline_elements&lt;/a&gt; themselves are [&lt;code&gt;Word of string], [&lt;/code&gt;Space], [`Code_span _] and so on. That means &lt;em&gt;each word&lt;/em&gt; and &lt;em&gt;each space&lt;/em&gt; has its own location recorded in the AST. This seemed a big waste, since we never (ish) use them, so the third optimisation was just to join the words and spaces up into one entry. With these three in place, the space usage went down by 6-7 times to a much more reasonable size.&lt;/p&gt;
&lt;p&gt;While in an upstreaming mood I&apos;ve also pulled out several other optimisations that were made for handling OxCaml-produced heavily templated code.&lt;/p&gt;
&lt;h3&gt;Odd&lt;/h3&gt;
&lt;p&gt;Odd is my tool to simulate the &amp;quot;First Class Docs&amp;quot; world by installing odoc files with &lt;code&gt;odd_driver&lt;/code&gt; (a mini fork of &lt;code&gt;odoc_driver&lt;/code&gt;) as an opam hook. I thought I might try getting an AI agent to use it to help write some OCaml programs, and then it could critique the tool and suggest improvements. This turned out to be very useful. I set up a &lt;a href=&quot;https://tangled.org/jon.recoil.org/odd/commit/552ba50ce629d4d7cb9a889088879c0a7f7fa963&quot;&gt;few tasks&lt;/a&gt; that required using the Jane Street libraries, as these tend to have mli files that are less easily read than many others in the OCaml ecosystem. Not only did it end up suggesting a number of useful bugfixes and improvements to &lt;code&gt;odd&lt;/code&gt;, but it also found several bugs in odoc that I submitted fixes for. A fix for &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1482&quot;&gt;extended opens&lt;/a&gt;, a fix to &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1484&quot;&gt;prevent hidden items leaking into the docs&lt;/a&gt;, some &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1485&quot;&gt;inconsistent strengthening of module types&lt;/a&gt;. It also spurred me on to fixing a long-standing issue - how to handle a hand-written wrapper module that &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1483&quot;&gt;doesn&apos;t expose everything&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I was using a mix of Claude and &lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731&quot;&gt;DeepSeek-V4-Flash-0731&lt;/a&gt; for these tests - the latter running entirely locally on my laptop with &lt;a href=&quot;https://github.com/antirez/ds4&quot;&gt;antirez/ds4&lt;/a&gt;. For the agent for DeepSeek I used &lt;a href=&quot;https://pi.dev/&quot;&gt;pi.dev&lt;/a&gt;. While being slower than Claude, this model was perfectly capable of passing each and every test, which I found quite amazing.&lt;/p&gt;
&lt;p&gt;The obvious next step was to see whether &lt;code&gt;odd&lt;/code&gt; was actually helpful to the agents. To this end I changed the tasks to have 3 different prompts - one with just the task, one suggesting to look in the opam lib dir for the mli files, and one suggesting using &lt;code&gt;odd&lt;/code&gt;. I set these all going, using openrouter rather than ds4 for speed. I&apos;ve not completely analysed the results yet but it does seem that &lt;code&gt;odd&lt;/code&gt; &lt;em&gt;isn&apos;t&lt;/em&gt; significantly helpful!&lt;/p&gt;
&lt;h2&gt;Docs CI&lt;/h2&gt;
&lt;p&gt;I switched over ocaml.org to the new docs CI running on &lt;a href=&quot;https://dill.caelum.ci.dev/&quot;&gt;dill.caelum.ci.dev&lt;/a&gt;. This is using &amp;quot;day11&amp;quot;, which is a variant of &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark Elvers&lt;/a&gt;&apos;s &lt;a href=&quot;https://github.com/mtelvers/day10&quot;&gt;day10&lt;/a&gt;. The switch-over was relatively smooth, but there were two issues. The first was that the rendering of package-wide markdown docs disappeared, leading to a &lt;a href=&quot;https://github.com/ocaml/ocaml.org/pull/3731&quot;&gt;PR to revert the switch&lt;/a&gt;. Fortunately I was able to &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci/commit/83097578dff71936913c43f2d0cbafd73affdc05&quot;&gt;fix it&lt;/a&gt; and rebuild everything within a day, so the PR was closed without merging.&lt;/p&gt;
&lt;p&gt;The second issue was a bit more insidious. The docs CI machine, dill, has a layered storage solution - a pair of mirrored 7TB spinning disks with a 2TB SSD dm-cache in front. At that point I was doing manual pruning of the old layers, and it had got a bit full. I triggered a cache prune at roughly the same time as a big rebuild came in (within an hour or so, at least), and this ended up completely filling up the 2TB dm-cache disk. This caused the whole machine to grind to a halt for a few hours.&lt;/p&gt;
&lt;p&gt;The way it&apos;s &lt;em&gt;supposed&lt;/em&gt; to work is that when someone requests a documention page on ocaml.org, the ocaml.org server requests a json representation of the page from dill, then renders that as the body back to the client. In order to make this as fast as possible, it&apos;s not the ocaml-docs-ci server that actually sends the page back, but a caddy server sitting in front of it. This works nicely most of the time, but when the machine was busy with flushing the dm-cache blocks it was taking quite a long time to respond to these requests. This led to the ocaml.org server running out of fds and crashing!&lt;/p&gt;
&lt;p&gt;The two-fold solution was firstly to increase the fd limit on the ocaml.org server, and secondly to stick the layer pruning in a cron job that runs in the middle of the night. That keeps the size of the layers used in the docs down to around 1TB, and so we shouldn&apos;t run out of cache blocks.&lt;/p&gt;
&lt;h3&gt;Monitoring&lt;/h3&gt;
&lt;p&gt;The old docs-ci machine was quite difficult to maintain. It was hard to see what it was doing, why it was doing it, and to get to the relevant logs. As a consequence, the new incarnation had much more of a focus on being able to figure out what&apos;s going on. This manifests in two ways - firstly being able to browse the packages, the build and docs logs and classifying failures correctly, and secondly by having a much more useful grafana dashboard telling me how healthy the machine is.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;grafana.png&quot; alt=&quot;The Grafana dashboard&quot; &gt;&lt;/p&gt;
&lt;p&gt;The interesting thing about this dashboard is that Claude created it entirely, and via config files rather than driving the UI. It&apos;s not perfect - the &apos;Uptime&apos; graph is a little odd - but it certainly shows me almost everything that&apos;s interesting. The screenshot only shows the top of the dashboard - there&apos;s a lot more if you scroll down.&lt;/p&gt;
&lt;h2&gt;OxCaml VMM&lt;/h2&gt;
&lt;p&gt;A new project. I had a catchup with &lt;a href=&quot;https://github.com/djs55&quot;&gt;Dave Scott&lt;/a&gt; and while reminicing about the good ol&apos; XenServer days he mentioned the project &lt;a href=&quot;https://github.com/libkrun/libkrun&quot;&gt;libkrun&lt;/a&gt;. An OCaml version of this sounded like fun - and an OxCaml one even more so, as we should be able to use some of the OxCaml features to make it really quite efficient.&lt;/p&gt;
&lt;p&gt;It&apos;s very early days for this now, but it&apos;s capable of booting a linux VM on my macbook and on my Raspberry Pi 4. It&apos;s already a useful playground for looking at how modes can help us write idiomatic looking code that doesn&apos;t allocate. Right now there&apos;s nothing to be gained with the multicore modes, but that may happen as I move beyond one vCPU. I&apos;ll do a bit more of a write-up of this soon. Code is &lt;a href=&quot;https://tangled.org/jon.recoil.org/ocaml-microvm&quot;&gt;here&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;A bit of fun&lt;/h2&gt;
&lt;p&gt;While I was waiting in the hospital I needed something to take my mind off the surgery. I can&apos;t remember what got me on this track, but I was thinking about old Amiga demos, and I remember one in particular - &lt;a href=&quot;https://www.pouet.net/prod.php?which=1477&quot;&gt;ARTE by Sanity&lt;/a&gt; that was spectacular. On effect I enjoyed particularly was a texture mapped sphere. It was something I had tried to write myself back in the day, involving, if I recall correctly, a lookup table that was created using AMOS basic. I wondered if Claude could take a look into it and figure out &lt;em&gt;precisely&lt;/em&gt; how it worked.&lt;/p&gt;
&lt;p&gt;It was fascinating watching Claude work at the problem. It pulled down a copy of &lt;a href=&quot;https://github.com/BlitterStudio/amiberry&quot;&gt;amiberry&lt;/a&gt;, which exposes a socket over which you can control the emulator. I had bought a copy of &lt;a href=&quot;https://www.amigaforever.com/&quot;&gt;Amiga Forever&lt;/a&gt; a few years back so I had all the kickstart roms and workbench disks around, so it used them, downloaded the ARTE adf and booted it up. I then paused it at the effect I was interested in, and Claude interrogated the state of the VM over the socket. It then searched around and found a copy of the &lt;a href=&quot;https://aminet.net/dev/asm/SanityOperationSystem.readme&quot;&gt;Sanity Operation System&lt;/a&gt; which matched was was used to assemble the demo, in particular figuring out where local state was stored. It took quite a while to figure out the bitplane trickery, and there were several round trips where the output was messed up in various ways, but it got there in the end.&lt;/p&gt;
&lt;p&gt;Of course, most of the interesting bits about the effect is that it worked at all on the Amiga hardware, and we miss out on all that in reproducing in modern OCaml, but nevertheless it&apos;s quite fun to see it running!&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;arte.png&quot; alt=&quot;Arte Demo&quot; &gt;&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/09/summernotes.html" rel="alternate" title="Summer notes"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/06/first-class-docs.html</id>
    <title type="text">First Class Docs in OCaml</title>
    <updated>2026-06-22T00:00:00Z</updated>
    <published>2026-06-22T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;The quality of reference documentation available in OCaml for libraries has improved enormously over the past few years. Docs are now built for every package in opam, hosted on ocaml.org, with search, source rendering, inline media, a navigation sidebar and many other features.&lt;/p&gt;
&lt;p&gt;However, it&apos;s still not the case that documentation in OCaml has what I&apos;d call &amp;quot;First Class&amp;quot; status. What do I mean by this? Well, let&apos;s give some examples of what I think it &lt;em&gt;should&lt;/em&gt; look like.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;High quality API and package documentation should be installed with every opam package, meaning that all build systems ought to be able to do this, and that docs should be checked by CI systems.&lt;/li&gt;
&lt;li&gt;Editors should offer completion in docs, highlight any errors when you&apos;re editing with clear actions to fix them, and display rendered docs natively.&lt;/li&gt;
&lt;li&gt;It should be possible to read and search through the documentation easily, through your IDE and in the terminal.
So why aren&apos;t we there yet? What needs to be done? Let&apos;s first start by explaining a little about why it&apos;s harder in OCaml than in many other languages. If you already know this, skip to the &lt;a href=&quot;./#what_to_do&quot;&gt;what to do about it&lt;/a&gt; section!&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Why it&apos;s tricky&lt;/h2&gt;
&lt;p&gt;API docs in OCaml are written using specially formatted comments, much like they are in &lt;a href=&quot;https://www.doxygen.nl/&quot;&gt;many&lt;/a&gt; &lt;a href=&quot;https://haskell-haddock.readthedocs.io/latest/markup.html&quot;&gt;other&lt;/a&gt; &lt;a href=&quot;https://doc.rust-lang.org/rust-by-example/meta/doc.html&quot;&gt;languages&lt;/a&gt;. We usually write our documentation in &lt;code&gt;mli&lt;/code&gt; files:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;type t
(** The docs for the type [t] go here *)

val wrangle : t -&amp;gt; t list
(** Here we can talk about [wrangle] and how it relates to {!t} *)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;What makes OCaml more tricky than many other languages is OCaml&apos;s module system. A simple example is this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;type t
(** Here are my docs for [t] *)

module TMap : Map.S with type t = t
(** This module is a map from the standard library *)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;module &lt;code&gt;TMap&lt;/code&gt; is a module found within &lt;em&gt;our&lt;/em&gt; module, containing lots of types and values that all need docs, but its definition is in &lt;em&gt;another&lt;/em&gt; module - in the standard library in this case. The standard library isn&apos;t particularly special, and functors aren&apos;t the only way this can happen, so to produce the correct docs we need to have access to the docs of &lt;em&gt;all&lt;/em&gt; of our dependencies, just in case they turn out to be necessary to fully document our own interfaces.&lt;/p&gt;
&lt;p&gt;The information we need is all in the typedtrees - the &lt;code&gt;.cmt&lt;/code&gt; and &lt;code&gt;.cmti&lt;/code&gt; files - of our own files and all our dependencies, but there&apos;s a non trivial amount of work that goes into processing these, so we effectively &lt;em&gt;compile&lt;/em&gt; those &lt;code&gt;.cmti&lt;/code&gt; files into &lt;code&gt;.odoc&lt;/code&gt; files to make this process efficient.&lt;/p&gt;
&lt;p&gt;The upshot of this is that in order to produce fully processed docs for any particular package, we need to have the &lt;code&gt;.odoc&lt;/code&gt; files for all of the dependency packages available somewhere.&lt;/p&gt;
&lt;p&gt;The approaches that work today all build the docs &lt;em&gt;after&lt;/em&gt; the package has been installed, which is how tools like like &lt;a href=&quot;https://erratique.ch/software/odig&quot;&gt;odig&lt;/a&gt; and &lt;a href=&quot;/reference/nox-odoc-driver/&quot;&gt;odoc_driver&lt;/a&gt; work. But what would it take to ensure it happens at package install package, and for &lt;em&gt;all&lt;/em&gt; packages?&lt;/p&gt;
&lt;h2&gt;What to do?&lt;/h2&gt;
&lt;p&gt;Odig exists, of course, and it does a great job of building docs when you ask for them. But it&apos;s an optional install, and the odoc files it builds are squirreled away in its own &lt;code&gt;var/cache/odig&lt;/code&gt; directory, and not necessarily up to date, by design.&lt;/p&gt;
&lt;p&gt;The key thing is to define the standard place in the opam tree where the odoc and odocl files will live, and ensure that packages are responsible for doing this. With this in place, we simply know that, if a package is installed, its libraries will have their compiled module docs installed and its package docs will be installed in their appropriate places. This gives all other tooling a reliable source of &lt;code&gt;.odoc&lt;/code&gt;/&lt;code&gt;.odocl&lt;/code&gt; files for whatever they need them for.&lt;/p&gt;
&lt;h3&gt;Build system and CI integration&lt;/h3&gt;
&lt;p&gt;The most-used developer tool to build docs today is dune, and it&apos;s woefully short of where we&apos;d like it to be. Without the odoc files from other packages the docs it builds are full of unresolved links and unexpanded modules, and the error messages it gives are almost useless.&lt;/p&gt;
&lt;p&gt;We&apos;re working on a &amp;quot;local&amp;quot; fix for this, meaning that dune will know how to build docs for the dependency packages. The docs produced are far better, the error messages much more useful and it will greatly improve the standard of self-published docsets and what&apos;s published on ocaml.org. However, it would be far better if the docs for the dependency packages were installed already. Then dune would only have to concern itself with the docs of the workspace packages, and the implementation would be simpler.&lt;/p&gt;
&lt;p&gt;For a truly first-class docs experience, &lt;em&gt;all&lt;/em&gt; OCaml build systems would have to be able to build docs themselves. We&apos;ve got quite clear instructions on &lt;a href=&quot;/reference/nox-odoc/driver.html&quot;&gt;how to drive odoc&lt;/a&gt; at a low level, but we need better instructions at a higher level, for which the main source of info right now is the reference driver.&lt;/p&gt;
&lt;p&gt;As a stop-gap measure, we could tweak &lt;a href=&quot;/reference/nox-odoc-driver/&quot;&gt;odoc_driver&lt;/a&gt; and use it as a black box for doing the job. We can use this for all packages that don&apos;t yet install their own docs as a post-processing step after the opam build. Clearly this would be significantly less developer friendly than integrating it properly with the build system. For example, it wouldn&apos;t be able to do incremental rebuilds.&lt;/p&gt;
&lt;p&gt;Once we&apos;ve got the build systems working, checking things with CI would be pretty straightforward. Even before we&apos;re at that point, we could wire up &lt;a href=&quot;/reference/nox-odoc-driver/&quot;&gt;odoc_driver&lt;/a&gt; which already provides reasonable error messages.&lt;/p&gt;
&lt;h3&gt;Editor support&lt;/h3&gt;
&lt;p&gt;The main additional features I&apos;d like to see in editors would be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Tab-completion of references. With all the odoc files installed this would be easy enough, but I&apos;m not sure if it would be quick enough right now. If it isn&apos;t, it should be straightforward enough to add in an index.&lt;/li&gt;
&lt;li&gt;Error reporting. This is a bit trickier, but is not dissimilar from how Merlin does its work - we would rely on the build system to build the odoc files for the rest of the project, then have something that can report errors from a buffer.&lt;/li&gt;
&lt;li&gt;Preview of docs. When incremental builds of docs are available, this should be straightforward. The dune rules we&apos;ve been working on already can do this.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Search&lt;/h3&gt;
&lt;p&gt;And, of course, there&apos;s search. We currently have search in odoc output powered by &lt;a href=&quot;/reference/nox-sherlodoc/&quot;&gt;Sherlodoc&lt;/a&gt;. However, Sherlodoc is focused on searching of API functions, and doesn&apos;t index mld pages at all. We can augment sherlodoc with search powered by other mechanisms, such as BM25 or using LLM embeddings, as we demonstrated &lt;a href=&quot;/blog/2025/08/ocaml-mcp-server.html&quot;&gt;last year&lt;/a&gt;, to provide a more holistic search function for package documentation.&lt;/p&gt;
&lt;p&gt;So now let&apos;s focus on how we get to the point where all the odoc files are installed.&lt;/p&gt;
&lt;h2&gt;The path to success&lt;/h2&gt;
&lt;h3&gt;What gets installed?&lt;/h3&gt;
&lt;p&gt;First we need to define what actually gets put into the filesystem. The docs we&apos;ve got on driving odoc are great at the level of individual commands, but not anything with a ecosystem-wide view on actually how to lay out the files, where to put package docs vs library docs vs rendered source files, what we definitely need installed, and what&apos;s more of a convenience file that&apos;s generated while we&apos;re building docs.&lt;/p&gt;
&lt;p&gt;The starting point for this is probably what &lt;a href=&quot;/reference/nox-odoc-driver/&quot;&gt;odoc_driver&lt;/a&gt; produces as part of the builds for ocaml.org. We can easily see what these are by just running &lt;code&gt;odoc_driver&lt;/code&gt; and supplying output paths for the intermediate files. For example, &lt;code&gt;odoc_driver --odoc-dir _odoc --odocl-dir _odocl&lt;/code&gt;. These files are more-or-less what should be installed somewhere in your opam switch directory.&lt;/p&gt;
&lt;p&gt;Having figured all this out, we then need to figure out how to make that happen. Let&apos;s first consider the ambitious plan - one that might take some time to do.&lt;/p&gt;
&lt;h3&gt;The ambitious plan&lt;/h3&gt;
&lt;p&gt;Let&apos;s upstream the necessary parts of odoc into &lt;a href=&quot;https://github.com/ocaml/ocaml&quot;&gt;ocaml/ocaml&lt;/a&gt;!&lt;/p&gt;
&lt;p&gt;If we want packages to install their own docs, and we need odoc available to do that, then we should arrange for odoc to be part of the first package that needs docs - and that would be the compiler itself. Otherwise, the way we work with stdlib will always be different from many other packages.&lt;/p&gt;
&lt;p&gt;We don&apos;t need &lt;em&gt;all&lt;/em&gt; of odoc, just the parts that are essential for other packages to be able to produce &lt;em&gt;their&lt;/em&gt; docs correctly. So essentially we need enough to run the compile and link phases, and at least one output format.&lt;/p&gt;
&lt;p&gt;All of the tooling for additional uses - search, html output, completion, and so on, can be packaged up in additional tools packages. It&apos;s not completely clear exactly where to draw the line, so we&apos;ll have to figure that out as a community.&lt;/p&gt;
&lt;p&gt;The ocaml repository already contains &lt;a href=&quot;https://ocaml.org/manual/5.5/ocamldoc.html&quot;&gt;ocamldoc&lt;/a&gt; which is used to build the manual for OCaml. We&apos;ve had support for building the manual with odoc since &lt;a href=&quot;https://github.com/ocaml/ocaml/pull/9997&quot;&gt;OCaml 4.13&lt;/a&gt;, so replacing ocamldoc with odoc isn&apos;t too far of a stretch, particularly as odoc has really been the standard doc tool for OCaml for many years now.&lt;/p&gt;
&lt;h3&gt;The additional quicker plan&lt;/h3&gt;
&lt;p&gt;The ambitious plan will take a while to sort out, but while we&apos;re working on it we can explore what life would be like with this all working using a shortcut. If we have a tool that can install the odoc files on behalf of the packages without having to actually modify the packages to do it, we can work on some of the integration tasks and try them out.&lt;/p&gt;
&lt;p&gt;Odig obviously does a lot of this already, but it&apos;s missing quite a few of the modern odoc 3 features. Odoc_driver is closer to what we&apos;d need, but we need a couple of tweaks to make it really fit.&lt;/p&gt;
&lt;p&gt;We can then execute this package-by-package after they&apos;re installed, and if we do this using opam&apos;s post-install hooks, we can simulate what it would be like had the packages installed the docs themselves. It doesn&apos;t &lt;em&gt;quite&lt;/em&gt; work, an obvious example of which is that opam&apos;s view of what the packages have installed won&apos;t contain the odoc files. But it&apos;s certainly good enough to try out things like CLI-based doc viewing, or completion of references.&lt;/p&gt;
&lt;p&gt;With these changes, the &amp;quot;energy barrier&amp;quot; to writing good docs for your OCaml packages will be significantly reduced, and the ubiquity of the docs in all opam switches will result in tools and workflows that really make the most of this enormously useful, but currently underused resource.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/06/first-class-docs.html" rel="alternate" title="First Class Docs in OCaml"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/06/weeknotes-25.html</id>
    <title type="text">Weeknotes 2026 week 24-25</title>
    <updated>2026-06-22T00:00:00Z</updated>
    <published>2026-06-22T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;I have now finally finished almost all of the end-of-term duties, incuding marking of 128 Foundations of Computer Science exam questions. Phew! It ended up being quite a week, with the unwelcome additional burden of having my car written off due to a minor prang.&lt;/p&gt;
&lt;p&gt;The main things I&apos;ve been working on have been finalising the new ocaml-docs-ci implementation, and working on a plan for &amp;quot;first class docs&amp;quot; in OCaml. The latter I&apos;ve written up &lt;a href=&quot;/blog/2026/06/first-class-docs.html&quot;&gt;here&lt;/a&gt;. The former is currently running on a rebuilt &lt;a href=&quot;https://dill.caelum.ci.dev&quot;&gt;https://dill.caelum.ci.dev&lt;/a&gt;, building docs for all versions of all packages.&lt;/p&gt;
&lt;h3&gt;First-class docs&lt;/h3&gt;
&lt;p&gt;I&apos;ve written up a post on what &apos;First Class Docs&apos; in OCaml might mean, and in parallel I&apos;ve made a little tool to see how it feels. I&apos;ve not put it on the main post, but I&apos;ll share a little video of it here:&lt;/p&gt;
&lt;p&gt;&lt;video controls src=&quot;./odd_demo.mp4&quot;&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;It&apos;s a tiny tool that auto-builds docs for your switch, keeping them up to date automatically as packages are added and removed. It&apos;s got sherlodoc-based search, it can show docs in markdown in the terminal for any item in any package, completion of odoc-style references, and the completion is integrated with zsh so you can do `odd doc Odoc_&amp;lt;tab&amp;gt;` and it will show you sensible completions.&lt;/p&gt;
&lt;p&gt;Next step is to create a skill for LLMs to use this tool. I&apos;m pretty sure they&apos;ll find it very helpful!&lt;/p&gt;
&lt;p&gt;Source is here: &lt;a href=&quot;https://tangled.org/jon.recoil.org/odd&quot;&gt;https://tangled.org/jon.recoil.org/odd&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;Docs CI&lt;/h3&gt;
&lt;p&gt;The work on docs-ci was not terribly exciting, but should be leading to a switch over in the next few weeks, once I&apos;m happy that the service is more stable than the current one (which isn&apos;t a particularly high bar).&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The new way of running meant fewer docker containers, so the deployment had to be changed&lt;/li&gt;
&lt;li&gt;I added a Caddy webserver rather than nginx, and updated the configuration so that it can serve multiple profiles&lt;/li&gt;
&lt;li&gt;I&apos;ve only enabled the &apos;quick&apos; profile and the &apos;full&apos; profile for now - the other two I&apos;ve been using are &apos;oxcaml&apos; and &apos;odoc-master&apos; that do the obvious things.&lt;/li&gt;
&lt;li&gt;I&apos;ve added a few more bits of info when you&apos;re browsing the state - things like the &lt;a href=&quot;https://dill.caelum.ci.dev/profiles/full/snapshots/2ff3cda8c7e6&quot;&gt;latest commit&lt;/a&gt; on the opam repositories being used, a &lt;a href=&quot;https://dill.caelum.ci.dev/profiles/full/snapshots/a315b93bb320/diff/eabbc6b9a9da&quot;&gt;diff view&lt;/a&gt; between snapshots, some instructions on &lt;a href=&quot;https://dill.caelum.ci.dev/profiles/full/p/raylib/2.1.0&quot;&gt;what to do when your package fails&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Some more minor tidying for &amp;quot;release&amp;quot;
I&apos;ve now left it running for a while just to see how it behaves. It&apos;s looking pretty good, though there were a few issues to iron out. The first was to do with epochs, which is a mechanism to allow deployment of new versions of odoc, sherlodoc, odoc_driver and so on whilst keeping the older version &amp;quot;live&amp;quot; until you want to switch over to the new version - the new epoch. The problem was that it was keying the epoch off the full solves of the tools rather than just the versions of the tools themselves. The consequence was far too rapid cycling of epochs, so I had to fix that. The second was that the switch to the new epoch did a synchronous GC of the older epochs, which led to a very long pause. Then the release of OCaml 5.5 turned out to be very useful, as while all the packages built as expected, none of the docs appeared. We have an override in the profile to be able to select the compiler version used to build odoc_driver and related tools, which was set to &apos;None&apos; to mean &apos;latest version&apos;. However, in this case, &apos;latest version&apos; meant &apos;pin to the latest version of the compiler&apos; as opposed to &apos;best version that has a solution&apos;. This constraint, along with js_of_ocaml not yet working with OCaml 5.5, meant that there was no solution for the tools, hence no docs! The quick fix for this was to put a pin in place to OCaml 5.4.1 and the docs popped right out.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Some info: A full build took about 12 hours. This is on roughly the same hardware as the current docs ci, where it takes the best part of a week to do the same, so the day10/day11 architecture has made it much faster. Each snapshot builds about 17600 packages, in about 31,000 build layers, and 31,000 doc layers. The total space for one complete snapshot is on the order of 1.1Tb. The OCaml 5.5 release rebuilt most of the packages, so we&apos;re currenly up to about 2.3Tb of storage used.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/06/weeknotes-25.html" rel="alternate" title="Weeknotes 2026 week 24-25"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/06/weeknotes-23.html</id>
    <title type="text">Week 23, June 2026</title>
    <updated>2026-06-08T00:00:00Z</updated>
    <published>2026-06-08T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;...and previous weeks.&lt;/p&gt;
&lt;h2&gt;Current topics&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;./#general&quot;&gt;General odoc directions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;./#ocaml_docs_ci&quot;&gt;OCaml Docs CI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;./#dune_odoc_rules&quot;&gt;Dune Odoc rules&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;./#compsci&quot;&gt;Exams!&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;General odoc directions&lt;/h2&gt;
&lt;p&gt;While docs are, in general, better than they were a few years back, it&apos;s definitely not &amp;quot;done&amp;quot;. I&apos;ve been considering what life would look like if docs were &amp;quot;first class&amp;quot;. I&apos;m writing this up, but for now I&apos;ll at least point at the first few bullet points.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;High quality API and package documentation should be installed with every opam package. This implies that every build system is capable of producing them, and CI should check them.&lt;/li&gt;
&lt;li&gt;Editors should offer completion in docs, highlight errors when you&apos;re editing, and offer a way to display docs natively.&lt;/li&gt;
&lt;li&gt;Searching for docs, in editor or via the CLI, should be straightforward and useful.
The last point is particularly useful for LLMs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Obviously there&apos;s a lot of work required here! I&apos;m currently thinking about a pathway that gets us most of the benefits ASAP, with a longer term outlook of doing it without &amp;quot;hacks&amp;quot;. I&apos;d like to get the &apos;ASAP&apos; tooling published so that other people can have a go and get some of these benefits.&lt;/p&gt;
&lt;h3&gt;.cmt vs .odoc&lt;/h3&gt;
&lt;p&gt;As part of this, I wanted to take a good look at the cmt/cmti format to see what&apos;s in there and why, and where the odoc file format differs. They both are based on the underlying source tree, so it&apos;s quite reasonable to ask where the similarities and differences are, especially if I&apos;m suggesting installing odoc files as well as cmti!&lt;/p&gt;
&lt;p&gt;I realised in doing this that while we&apos;ve got docs for &lt;em&gt;users&lt;/em&gt; of odoc, we&apos;ve not got great coverage for &lt;em&gt;developers&lt;/em&gt;. So I&apos;ve spent some time this week working on this. I initially started writing this for this blog, but quickly realised that it really should be in the odoc repo, so that&apos;ll be coming soon. I&apos;ll try to get this done without spending too much time polishing it.&lt;/p&gt;
&lt;h3&gt;Idents vs Identifiers&lt;/h3&gt;
&lt;p&gt;In the interests of seeing closer parallels between odoc files and cmt/i files, I got Claude to spend some time investigating using local idents instead of identifiers in the odoc files, and switching to identifiers when writing out the odocl files. This would remove one big source of difference between the formats, and potentially make the code more easier to read for people used to OCaml&apos;s module implementation. A side benefit could also be an improvement in performance, as we spend quite a lot of time converting between these two types as odoc runs. Unfortunately this proved a little too tricky for Claude and while it made some progress, it seems I&apos;ll need to spend a lot more time describing the change, and/or working through it with more supervision.&lt;/p&gt;
&lt;h2&gt;OCaml Docs CI&lt;/h2&gt;
&lt;p&gt;This has been dragging. So I&apos;m currently on a sprint to get this sorted by cutting out any nice-to-haves. It&apos;s already in a state where it&apos;s better than current ocaml-docs-ci, so my focus now is on doing the absolutely necessary bits to have it in a sensible state for ocaml.org. These are:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Deployability - fix up the docker files / ansible / ocurrent-deployer such that getting a new version running is as simple as pushing to github.&lt;/li&gt;
&lt;li&gt;Wire in the epoch management mechanism. This has proven invaluable with the current ocaml-docs-ci, and was relatively simple to put in place.&lt;/li&gt;
&lt;li&gt;Ensure docs are good enough for others to debug/maintain/fix it.
So the other neat bits of it are deprioritised - building oxcaml packages, building latest odoc master branch, local override repositories, etc. Let&apos;s get it out and being used so it&apos;s no longer dragging me down!&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Dune Odoc rules&lt;/h2&gt;
&lt;p&gt;Paul-Elliot and Arthur have been doing a good job cutting up my branch into smaller chunks, but it seems to have hit a got a bit mired in the mud. There are a few good small commits, and a few have been upstreamed, but there&apos;s going to be a big chunk that&apos;s still quite hard to put into manageable commits. I&apos;ll be looking over what&apos;s been done so far to see if I can figure out a pathway forward.&lt;/p&gt;
&lt;h2&gt;Exams!&lt;/h2&gt;
&lt;p&gt;The exam for the Foundations of Computer Science course I lectured this year is coming up this week, so I&apos;ll be spending a lot of time marking for a while!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/06/weeknotes-23.html" rel="alternate" title="Week 23, June 2026"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/05/weeknotes-18-19.html</id>
    <title type="text">Weeknotes May 2026 weeks 18-19</title>
    <updated>2026-05-11T00:00:00Z</updated>
    <published>2026-05-11T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Over the past two weeks I&apos;ve been mainly wrestling with odoc.&lt;/p&gt;
&lt;p&gt;I made a release of odoc - &lt;a href=&quot;https://github.com/ocaml/odoc/releases/3.2.0/&quot;&gt;3.2.0&lt;/a&gt; - and a &lt;a href=&quot;https://github.com/ocaml/opam-repository/pull/29834&quot;&gt;PR to opam-repository&lt;/a&gt;. In parallel, I also just managed to make an ocaml-docs-ci job that tracked the master branch of odoc. I was pretty horrified therefore to see a glaring error in the docs build of &lt;code&gt;base&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;It turned out that one of the patches we made for OCaml 5.5.0, to support modular explicits, had an issue. It caught me by surprise because I didn&apos;t think that anything already existed that uses this, but in fact the new constructors in the AST are being used by pre-existing language constructs, so even without source changes we were exercising the new logic. This turned out to be quite useful as it showed that there was some missing transformations in the code causing a hard failure. Fortunately the fix was .&lt;/p&gt;
&lt;p&gt;With this success of the new docs ci, I thought we could do some more serious runs, so I set a job going building docs for the latest version of everything in opam using the master branch of odoc. After getting this running the following few days was spent sorting out the issues that came up -  and .&lt;/p&gt;
&lt;p&gt;The second here showed up a problem in the fix for . In that, as I &lt;a href=&quot;/blog/2025/09/odoc-bugs.html&quot;&gt;wrote about previously&lt;/a&gt;, the fix was to rewrite module types being included without inline module substitutions. The code that did that effectively rewrote the following:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;sig
  ...
  module Foo := Bar
  ...
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;as&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;sig
  module Foo = Bar
end with module Foo := Bar
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The problem that the docs ci found was an instance that looked more like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;sig
  module A
  ...
  module Foo := A
  ...
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and of course, in this, we can&apos;t simply do the above transformation as module &lt;code&gt;A&lt;/code&gt; is only defined within the scope of the signature, so we can&apos;t reference it from outside.&lt;/p&gt;
&lt;p&gt;A few months back I had rewritten this code so that rather than going through the &lt;a href=&quot;/reference/odoc/odoc.xref2/Odoc_xref2/Tools/index.html#fragmap-functions&quot;&gt;Fragmap functions&lt;/a&gt; it directly used a &lt;a href=&quot;/reference/odoc/odoc.xref2/Odoc_xref2/Subst/index.html&quot;&gt;substitution&lt;/a&gt;, but I hadn&apos;t quite completed this. It seemed a good opportunity to dig it out and give it a whirl, resulting in a &lt;a href=&quot;https://github.com/jonludlam/odoc/commit/b9791be89ffac8f947eac880766e8d88147cb74e&quot;&gt;nice patch&lt;/a&gt;. With the power of the new docs CI I set this going, and while it fixed &lt;em&gt;that&lt;/em&gt; issue, it exposed another. At this point I felt that it was getting a bit hairy for a last-minute fix, as we needed to release odoc rather quickly for OCaml 5.5 support. So I shelved that plan, and instead went for the rather less dangerous fallback of doing nothing when we hit that pattern. The thought was that issue  is relatively rare, but the fragmap fix is applied to &lt;em&gt;all&lt;/em&gt; includes. The &apos;local substitution&apos; pattern is &lt;em&gt;also&lt;/em&gt; rare, as it was only a few packages that showed the error when I did the full docs run. So the chances of hitting &lt;em&gt;both&lt;/em&gt; we can expect to be even rarer! We went with  in the end.&lt;/p&gt;
&lt;p&gt;The other small defect I found was an interesting one. There is an  with odoc related to building docs for &lt;a href=&quot;https://colibri.frama-c.com/&quot;&gt;colibri2&lt;/a&gt;. What I happened to notice with the docs ci build was that the docs build was &lt;em&gt;succeeding&lt;/em&gt;. After a bit of debugging I finally figured out that this was due to missing dependency libraries during the build. Without these, it failed to do some expansions, which meant that it didn&apos;t hit the problematic case. This led to , and colibri2 is back to failing.&lt;/p&gt;
&lt;p&gt;Thanks to &lt;a href=&quot;https://github.com/panglesd&quot;&gt;panglesd&lt;/a&gt;&apos;s quick review, these were merged, and we were finally able to release &lt;a href=&quot;https://github.com/ocaml/odoc/releases/3.2.1&quot;&gt;3.2.1&lt;/a&gt; and put it .&lt;/p&gt;
&lt;h3&gt;Docs CI and universes&lt;/h3&gt;
&lt;p&gt;Docs CI has always had this concept of consistent &amp;quot;universes&amp;quot; of packages which are layered together. As a simplified example, imagine one universe for each version of the OCaml compiler. Then there are several versions of &lt;code&gt;findlib&lt;/code&gt;, each of which can compile with each version of &lt;code&gt;ocaml&lt;/code&gt;, so now we have &lt;code&gt;n*m&lt;/code&gt; possible universes. Obviously this would lead to a combinatorial explosion if we wanted to have every possible universe, but in practise there&apos;s a huge amount of overlap, even when we&apos;re trying to build every version of every package. So we exploit this by having a triple that defines every build - the package name, the package version, and a hash representing the dependencies of the package.&lt;/p&gt;
&lt;p&gt;A core principle of these universes is that the docs produced are &lt;em&gt;also&lt;/em&gt; consistent. For example, if your package was built with OCaml 5.4.1, no matter where you click, you&apos;ll never end up on a package that was build with a different version of OCaml. Similarly for every other dependency of the package whose docs you are browsing. You will always stay within the dependency universe.&lt;/p&gt;
&lt;p&gt;When we introduced &amp;quot;forward links&amp;quot; in odoc 3.0, we retained this. These forward links are links from a package&apos;s docs that should go to a reverse dependency of the package. For example, odoc&apos;s docs link to odoc-driver and odig, as examples of tools that use odoc. This is a straightforward extension of docs generation if all involved packages are already installed, as odoc has had split compile and link phases since very early versions. Compilation must be done in strict dependency order, as when odoc processes a &lt;code&gt;cmt(i)&lt;/code&gt; files it requires the &lt;code&gt;odoc&lt;/code&gt; files of dependencies. In contrast, linking only looks at &lt;code&gt;odoc&lt;/code&gt; files, so once we&apos;ve computed all of these we can link in any order we wish.&lt;/p&gt;
&lt;p&gt;The approach taken in &lt;code&gt;ocaml-docs-ci&lt;/code&gt; is different. Because we want to avoid processing packages more than once, we process each package individually. If a package has no forward links, we can do both compilation and linking for that package in one go. If package &amp;quot;a&amp;quot; has forward links to package &amp;quot;b&amp;quot; that depends upon &amp;quot;a&amp;quot;, we have to:&lt;/p&gt;
&lt;p&gt;1. compile the odoc files for package &amp;quot;a&amp;quot;, 2. compile the odoc files for package &amp;quot;b&amp;quot;, ensuring that the odoc files of package &amp;quot;a&amp;quot; are available 3. link the odoc files for package &amp;quot;a&amp;quot;, ensuring that the odoc files of package &amp;quot;b&amp;quot; are available 4. link the odoc files for package &amp;quot;b&amp;quot;, ensuring that the odoc files of package &amp;quot;a&amp;quot; are available.&lt;/p&gt;
&lt;p&gt;This is obviously more complex than if there were no forward links, as then it would be just a two step process - compile and link &amp;quot;a&amp;quot;, then compile and link &amp;quot;b&amp;quot;. The only complixity is in ensuring that the odoc files for &amp;quot;a&amp;quot; are available when doing the compile and link for &amp;quot;b&amp;quot;.&lt;/p&gt;
&lt;p&gt;A further wrinkle is that your forward links might require packages with version constraints. For example, odoc might have links to documented features in odoc-driver that are only in recent versions. Even worse: there might be multiple versions of &amp;quot;b&amp;quot; that all use the same version of &amp;quot;a&amp;quot;, and here we hit against the principle outlined above; that no matter where you click you never leave your universe. If there are multiple versions of &amp;quot;b&amp;quot; that link to the same version of &amp;quot;a&amp;quot;, then when you click the reverse links, you must end up back on the same &amp;quot;b&amp;quot; you started with. That means there must be multiple different versions of the docs for &amp;quot;a&amp;quot;, even if they&apos;re all built with the same set of build dependencies!&lt;/p&gt;
&lt;p&gt;So if we&apos;re being as efficient as possible, we need to divorce the build universe from the docs universe. That way we only need to build the package &amp;quot;a&amp;quot; once, even if we build docs for it multiple times. Currently day11 &lt;em&gt;doesn&apos;t&lt;/em&gt; do this, and is inefficiently doing multiple redundant builds for some packages. This doens&apos;t affect the &lt;em&gt;output&lt;/em&gt; as it&apos;s the docs universes that are important there, but it does mean it&apos;s wasting resources.&lt;/p&gt;
&lt;h3&gt;oi docs&lt;/h3&gt;
&lt;p&gt;All of which brings me on to &lt;code&gt;oi&lt;/code&gt;. I had a very useful discussion with Anil on whether we need to bring all of this complexity into &lt;code&gt;oi&lt;/code&gt; to build docs there. When you&apos;re working on a package in &lt;code&gt;oi&lt;/code&gt;, whilst it manages many different universes, it will only instantiate one for your package. Therefore we don&apos;t need to worry so much about staying in the same universe - there can be only one.&lt;/p&gt;
&lt;p&gt;So he suggested doing &lt;em&gt;late binding&lt;/em&gt; of the forward links somehow, to enable building docs alongside the standard package build. The forward links would be unresolved, but we could arrange for the link to be resolved once you installed the package to be linked to. This could be done in a variety of ways. One thought I had was that I could use a branch I&apos;d been working on a while ago which simply enumerates the possible link targets in an odocl file. The thing that makes this achievable is that we&apos;re only talking about &lt;em&gt;cross-package&lt;/em&gt; references, so a lot of machinery that exists within odoc is unnecessary. We&apos;ll only allow canonical reference, and scoping isn&apos;t a problem as they will all be fully qualified references.&lt;/p&gt;
&lt;p&gt;I&apos;ll be working on making this work for the rest of this week.&lt;/p&gt;
&lt;h3&gt;Random ideas&lt;/h3&gt;
&lt;p&gt;Some random ideas I had:&lt;/p&gt;
&lt;h4&gt;Package API diff&lt;/h4&gt;
&lt;p&gt;It&apos;s always been a nice idea to use the data we generate in docs CI in more creative ways than just producing the docs. A fairly obvious idea is to generate diffs to show what&apos;s changed in between versions of the packages. However, this was always a bit problematic because your own package doesn&apos;t necessarily uniquely define your API signatures. This is, of course, the whole reason we produce multiple universes of docs. However, in odoc 3.1 we added a neat feature that was intended to help you by only showing errors when they exist in your own packages, which worked by keeping track on a signature level of the package from which the signature came. What I realised is that we can use this info to discard diffs between package version builds so that they only show what&apos;s changed in the package being queried, rather than differences due to your dependencies.&lt;/p&gt;
&lt;p&gt;A quick example. Odoc has several maps such as &lt;a href=&quot;/reference/odoc/odoc.xref2/Odoc_xref2/Component/ModuleMap/index.html&quot;&gt;Component.ModuleMap&lt;/a&gt; that are constructed from the standard library functor, and therefore the elements in them you can see at that link have various &apos;since&apos; markers - but those all refer to the OCaml version, not the odoc version, even though these are odoc&apos;s docs. If we&apos;re trying to find the difference between two versions of odoc, we might observe differences in these just because we happened to compile the two odoc&apos;s with different compilers. Using the info put in for the warnings would eliminate this.&lt;/p&gt;
&lt;h4&gt;Fuzzing odoc-parser&lt;/h4&gt;
&lt;p&gt;We ought to fuzz odoc-parser. There are some nice tests in there - mostly expect tests - but a proper comprehensive fuzz will definitely show up some issues, as it&apos;s a hand-rolled parser. We spent some time a while back building a menhir parser, but the error paths were very hard to get working as well as the current implementation handles them. Fuzzing could help strengthen the case for using the menhir parser.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://gazagnaire.org&quot;&gt;Thomas&lt;/a&gt; has a neat &lt;a href=&quot;/blog/2026/05/https:/tangled.org/gazagnaire.org/monopampam/blob/main/ocaml-claude-skills/plugins/monopam/skills/fuzz/SKILL.html&quot;&gt;fuzzing skill&lt;/a&gt; for Claude that he&apos;s been using to great effect with his &lt;a href=&quot;https://gazagnaire.org/blog/2026-04-15-ccsds-protocol-stack.html&quot;&gt;protocol development&lt;/a&gt; work, so I&apos;d quite like to take a look into using that to give that a whirl.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/05/weeknotes-18-19.html" rel="alternate" title="Weeknotes May 2026 weeks 18-19"/>
    <category term="odoc"/>
    <category term="day11"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/04/weeknotes-2026-16-17.html</id>
    <title type="text">Weeknotes 2026 weeks 16-17</title>
    <updated>2026-04-28T00:00:00Z</updated>
    <published>2026-04-28T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;A two week update this week. Most of this fortnight has been spent on different sides of the same problem: getting OCaml documentation into a state where an LLM (or a human) can actually rely on it — search, packaging, performance, and infrastructure.&lt;/p&gt;
&lt;h2&gt;Docs and LLMs&lt;/h2&gt;
&lt;p&gt;It seems fairly obvious, but having docs available for libraries is very important for LLMs to be able to use them effectively, especially if they&apos;re new or private. There have been various papers on the subject — for instance the &lt;a href=&quot;https://arxiv.org/abs/2207.05987&quot;&gt;2022 paper on tool documentation&lt;/a&gt; and the more recent &lt;a href=&quot;https://arxiv.org/abs/2503.15231&quot;&gt;work on the importance of example code&lt;/a&gt; — as well as our own &lt;a href=&quot;https://toao.com/blog/ai-existential-ocaml&quot;&gt;contribution&lt;/a&gt; to the OCaml workshop last year.&lt;/p&gt;
&lt;p&gt;Examples in OCaml libraries often live in &lt;a href=&quot;https://ocaml.org/docs/generating-documentation#generating-mld-documentation-pages&quot;&gt;mld files&lt;/a&gt;, and the correctness of these can be automatically tested using &lt;a href=&quot;https://github.com/realworldocaml/mdx&quot;&gt;mdx&lt;/a&gt;, which works well on mld files.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://doc.sherlocode.com&quot;&gt;Sherlodoc&lt;/a&gt; deliberately skips mld files, because plain English in mld pages drowns the API hits — when indexing was originally enabled it ended up polluting the API results too much, with generic text appearing way too often. Sherlodoc doesn&apos;t do any of the interesting stemming or other BM25-style weighting, so as more mld content lands in packages we&apos;ll likely need to investigate a hybrid approach to searching the docs.&lt;/p&gt;
&lt;h2&gt;Oi!&lt;/h2&gt;
&lt;p&gt;Anil&apos;s just &lt;a href=&quot;https://anil.recoil.org/notes/2026w16&quot;&gt;shipped&lt;/a&gt; a neat new tool in OCaml land — &lt;a href=&quot;https://github.com/avsm/oi&quot;&gt;oi!&lt;/a&gt;, built on top of &lt;a href=&quot;https://github.com/mtelvers/day10&quot;&gt;day10&lt;/a&gt; and &lt;a href=&quot;https://www.dra27.uk/blog/platform/2025/12/17/its-merged.html&quot;&gt;relocatable OCaml&lt;/a&gt;. This is a &lt;em&gt;really&lt;/em&gt; neat little tool that you can use to run OCaml tools and scripts without having a dedicated opam setup.&lt;/p&gt;
&lt;p&gt;As I&apos;ve been getting my Claude-built day10 successor, &lt;a href=&quot;./#day11&quot;&gt;day11&lt;/a&gt;, into shape as a replacement for the guts of the current ocaml-docs-ci, I thought I&apos;d check to see if the libraries I&apos;d made could work as a drop-in for the d10 part of oi (Anil&apos;s vendored copy of the relevant bits of day10). This indeed worked nicely, without huge impact on the rest of oi (aside from excising a chunk of it, of course).&lt;/p&gt;
&lt;p&gt;Next step is to see if we can get decent doc support into oi. It&apos;d be great to be able to run &lt;code&gt;oi doc&lt;/code&gt; or &lt;code&gt;oi doc search&lt;/code&gt; in a project and have accurate docs pop out. More on this to come!&lt;/p&gt;
&lt;h2&gt;Odoc performance investigation&lt;/h2&gt;
&lt;p&gt;A particular problem with getting the docs to &apos;just pop out&apos; is that some docs take a long time and too much memory to build with odoc. So much so that the Github Action workflow that generates them can just die, particularly when trying to build the docs for the oxcaml branches of base and core. So I&apos;ve spent a little time investigating this over the last couple of weeks, and made some quite significant progress.&lt;/p&gt;
&lt;p&gt;Headline figures on my laptop: wall time 549 s → 474 s (−14%), total allocation 381 GB → 230 GB (−40%), peak RSS 2.00 GB → 1.81 GB.&lt;/p&gt;
&lt;p&gt;Particularly bad for odoc was an include-expansion explosion: a single source line in container_intf.ml was being walked 10,777 times during html-generate, because each ppx_template monomorphisation produced its own Include whose expansion nested further Includes. I had previously noted this issue and tackled the most awful of the problems, but there was more to fix!&lt;/p&gt;
&lt;p&gt;The includes are effectively a workaround for the lack of layout polymorphism in the current oxcaml compiler. A quick investigation showed that of the 155,828 doc comments in container_intf.ml there were only 33 unique strings, once the ppx_template duplication was de-duped. Memoising the parsing of these led to a substantial improvement.&lt;/p&gt;
&lt;p&gt;In addition, many of the includes end up flattened in the resulting HTML due to their module-type being either a signature or pointing at a hidden item. Includes come with a fair bit of overhead, so rather than flattening them at the point we&apos;re generating the HTML, we can spot this pattern and flatten them much earlier in the process.&lt;/p&gt;
&lt;p&gt;The single biggest html-generate win was a one-liner. &lt;a href=&quot;https://github.com/ocaml/odoc/blob/master/src/html/link.ml#L9&quot;&gt;segment_to_string&lt;/a&gt; was doing &lt;code&gt;Format.asprintf &amp;quot;%a%s&amp;quot;&lt;/code&gt; where the &lt;code&gt;%a&lt;/code&gt; formatter just emits &lt;code&gt;&amp;quot;&amp;lt;kind&amp;gt;-&amp;quot;&lt;/code&gt; or nothing. A direct pattern-match killed ~60% of html-generate allocation on stdlib, ~50% on core.&lt;/p&gt;
&lt;p&gt;I&apos;ll tidy up these patches, double check that they don&apos;t affect the output, and see how much of an improvement they make on Anil&apos;s &lt;a href=&quot;https://github.com/avsm/oxmono&quot;&gt;oxmono&lt;/a&gt; repo.&lt;/p&gt;
&lt;h2&gt;Day11&lt;/h2&gt;
&lt;p&gt;I mentioned &lt;a href=&quot;/blog/2026/04/weeknotes-2026-15.html&quot;&gt;last time&lt;/a&gt; about replacing the innards of &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci&quot;&gt;ocaml-docs-ci&lt;/a&gt; with the day11 libraries. This has worked out really nicely. The shape of the resulting tool is that we&apos;ve got an ocurrent pipeline that watches &lt;a href=&quot;https://github.com/ocaml/opam-repository&quot;&gt;the opam-repository&lt;/a&gt; on GitHub, and triggers the day11 build when it notices a change. Each layer corresponds to an ocurrent job. On top of the normal UI I&apos;ve added in some pages that make it easier to spot emerging problems in the docs build. My ultimate plan is to put an LLM in charge of this so it can watch the status of it every day, and then let me know if it thinks I need to do something!&lt;/p&gt;
&lt;p&gt;Docs CI showing the status of the current &amp;quot;snapshot&amp;quot;&lt;/p&gt;
&lt;h2&gt;Improved site&lt;/h2&gt;
&lt;p&gt;I made some small improvements to the site too. I added a &lt;code&gt;@figure&lt;/code&gt; plugin, and a margin notes plugin &lt;code&gt;{&amp;amp;margin Margin notes look like this!}&lt;/code&gt; , and added tags to my pages. You&apos;re already reading the result — the Docs CI screenshot above uses the new &lt;code&gt;@figure&lt;/code&gt; plugin, and you can see the margin notes.&lt;/p&gt;
&lt;p&gt;Until next week!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/04/weeknotes-2026-16-17.html" rel="alternate" title="Weeknotes 2026 weeks 16-17"/>
    <category term="odoc"/>
    <category term="day11"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/04/weeknotes-2026-15.html</id>
    <title type="text">Weeknotes 2026 week 15</title>
    <updated>2026-04-14T00:00:00Z</updated>
    <published>2026-04-14T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Once again, the docs CI went down. This time, something had scribbled over the docker partition and so we needed to do a full build from scratch. Fortunately the docs themselves were not in a docker volume and so we didn&apos;t have to rebuild everything to get the HTTP server up and running for ocaml.org. However, we did have to set a full build going so that we can build docs for new packages.&lt;/p&gt;
&lt;p&gt;This keeps happening, and is very annoying! So, that brings me onto the next topic: day11.&lt;/p&gt;
&lt;h2&gt;Day11&lt;/h2&gt;
&lt;p&gt;I&apos;ve posted multiple times about &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark Elvers&apos;&lt;/a&gt; &lt;a href=&quot;https://github.com/mtelvers/day10&quot;&gt;day10&lt;/a&gt; project. For me, it was an obvious extension of this to get it to build docs, and, as with most of my work recently, I set Claude on the task. However, Claude failed to integrate it nicely into the codebase. Taking a closer look at day10, it&apos;s a specialised tool that does its thing really well, but isn&apos;t built in a way that makes it easy to adapt in the ways that the docs needed. Clearly separating the AI generated code from the hand-written code is very important, so rather than going that route, I decided that I&apos;d try to build a more general day10 with Claude - day11!&lt;/p&gt;
&lt;p&gt;It&apos;s already at the point where it&apos;s been able to build the docs for all packages in opam-repository relatively quickly. Running it on &lt;code&gt;dill&lt;/code&gt;, which is roughly equivalent to &lt;code&gt;sage&lt;/code&gt;, where the docs are currently built, it takes about 6 hours or so to build everything, where &lt;code&gt;sage&lt;/code&gt; with ocaml-docs-ci takes several days.&lt;/p&gt;
&lt;p&gt;Some intriguing directions we might take this in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It&apos;s a generic build plaform for opam packages, therefore could possibly be used for easily executing binaries from opam packages.&lt;/li&gt;
&lt;li&gt;It can build _itself_ - including new/different dependencies. Interesting for a self-modifying tool!&lt;/li&gt;
&lt;li&gt;It&apos;s easy to drop into a container with precisely the correct dependencies for any package, so useful for debugging build failures. This is partially implemented already.&lt;/li&gt;
&lt;li&gt;We can already provide overlay opam-repositories for testing of new/altered packages
One really nice test of whether the organisation of the libraries in day11 is correct is whether we can plumb it into the docs-ci ocurrent pipeline easily, and have the CLI tools for it coexist nicely.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Odoc performance&lt;/h2&gt;
&lt;p&gt;Running odoc on with some of the oxcaml libraries exposes some critical weaknesses of the current code - in particular &lt;a href=&quot;/blog/2026/03/weeknotes-2026-13.html#oxmono-docs-build&quot;&gt;performance problems&lt;/a&gt; with particular styles of code. We can&apos;t build the docs for Anil&apos;s &lt;a href=&quot;https://github.com/avsm/oxmono&quot;&gt;oxmono&lt;/a&gt; repo with GHA as it simply runs out of memory. I&apos;ve therefore been investigating some of the more egregious memory problems. I&apos;ve got quite a few patches already with some good results, but nothing yet that&apos;s going to make it into upstream odoc without some more testing.&lt;/p&gt;
&lt;h2&gt;Other bits and bobs&lt;/h2&gt;
&lt;p&gt;I had fun hour or so putting together an odoc plugin to replicate the experience of davesnx&apos;s &lt;a href=&quot;https://davesnx.github.io/parseff&quot;&gt;parseff site&lt;/a&gt;. The plugin is &lt;a href=&quot;https://tangled.org/jon.recoil.org/odoc-parseff-shell/&quot;&gt;here&lt;/a&gt;, and to use it, see my modified parseff repo:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-shell&quot;&gt;git clone https://github.com/jonludlam/parseff.git#odoc-plugins
opam switch create . --with-doc
dune build @doc
&lt;/code&gt;&lt;/pre&gt;
&lt;figure&gt;
  &lt;a href=&quot;https://jon.ludl.am/experiments/parseff&quot;&gt;&lt;img src=&quot;parseff.png&quot; alt=&quot;A screenshot of the parseff site&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;A screenshot of the parseff site produced by the plugin&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;There are a number of advantages and disadvantages to this. As @davesnx &lt;a href=&quot;https://sancho.dev/blog/ocaml-documentation-as-markdown&quot;&gt;wrote&lt;/a&gt;, his concern with the markdown output was to be able to integrate the odoc output seamlessly with an existing site, and it does this very well. However, it&apos;s at a cost - we lose links in the API docs, links to source, the source rendering itself, and so on. Whereas the plugin I made keeps all of those nice features, but is still tricky to integrate with a larger site.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/04/weeknotes-2026-15.html" rel="alternate" title="Weeknotes 2026 week 15"/>
    <category term="day11"/>
    <category term="odoc"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/04/odoc_and_ocaml_notebooks.html</id>
    <title type="text">Odoc and OCaml Notebooks</title>
    <updated>2026-04-06T00:00:00Z</updated>
    <published>2026-04-06T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;As the chief maintainer of OCaml&apos;s odoc, I&apos;m required to think hard about its future. What impact will advances in &lt;a href=&quot;https://anil.recoil.org/notes/aoah-2025&quot;&gt;agentic programming&lt;/a&gt;, &lt;a href=&quot;https://dl.acm.org/doi/10.1145/3759536.3763802&quot;&gt;collaborative literate coding&lt;/a&gt;, and the dramatic increase in web platform capabilities through &lt;a href=&quot;https://ocaml.org/tools/wasm-target&quot;&gt;wasm&lt;/a&gt; and WebGPU have — both on how odoc works and on how we develop it?&lt;/p&gt;
&lt;p&gt;I find the best way to explore these topics is to &lt;em&gt;build&lt;/em&gt; something, so I&apos;ve gone all-in on Claude and used it to rewrite my website with all new odoc-built dune-friendly client-side Jupyter-style notebooks as an integral part of it! Those of you that have read this blog for a while will know I&apos;ve been &lt;a href=&quot;/blog/2025/07/retrospective.html&quot;&gt;tinkering with notebooks&lt;/a&gt; for a year or so now, and I&apos;ve had an &lt;em&gt;extremely&lt;/em&gt; long-running fascination with &lt;a href=&quot;/blog/2025/12/an-svg-is-all-you-need.html&quot;&gt;scientific visualisation in the browser&lt;/a&gt;, but I really wanted to push this hard and see how adding Claude into the mix would work out.&lt;/p&gt;
&lt;figure&gt;
  &lt;a href=&quot;/notebooks/interactive_map.html&quot;&gt;&lt;img src=&quot;notebook.png&quot; alt=&quot;A client-side notebook using TESSERA&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;A client-side notebook using TESSERA&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;It&apos;s very helpful to have a specific use-case in mind for the notebooks to guide the exploratory work. This was an easy choice: My group has been spending a lot of time recently on &lt;a href=&quot;https://geotessera.org/&quot;&gt;TESSERA&lt;/a&gt;, and this provided a perfect demonstration for the notebooks, as it heavily relies upon being able to use interactive maps to choose areas of interest, and to visualise the embeddings and picture the output as map overlays. As well as being a fine motivating example, it&apos;s also given me an excellent reason to learn more about geospatial coding, which I&apos;ve been wanting to do for a while. The &lt;a href=&quot;https://github.com/ucam-eo/tessera-interactive-map&quot;&gt;specific example&lt;/a&gt; I chose to port required adding support for various interactive widgets in the browser-based programming environment, where I&apos;d only previously been able to render &lt;a href=&quot;/blog/2025/04/this-site.html&quot;&gt;static rich content&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Also, inspired by my colleagues&apos; &lt;a href=&quot;https://digitalflapjack.com/blog/marimo/&quot;&gt;experiences with notebooks&lt;/a&gt;, I&apos;ve been &lt;a href=&quot;./#execmodel&quot;&gt;experimenting with the execution model&lt;/a&gt; and interactivity to ensure reproducibility, and considering the implications of the FAIR (findability, accessibility, interoperability, reusability) principles of the &lt;a href=&quot;https://dl.acm.org/doi/10.1145/3759536.3763802&quot;&gt;Fairground paper&lt;/a&gt;, which provides a perfect framework of requirements for the odoc notebooks. These more ambitious goals meant a more ambitious approach to the changes.&lt;/p&gt;
&lt;p&gt;Despite this, with this rewrite I&apos;ve thought much harder to ensure that what I&apos;ve done will fit into the upstream ecosystem. It&apos;s now all built using dune with the standard rules I&apos;ve &lt;a href=&quot;/blog/2025/12/claude-and-dune.html&quot;&gt;already started upstreaming&lt;/a&gt;, and by using plugins with odoc I&apos;ve minimised the necessary changes to the core of odoc. While there&apos;s work to be done, there&apos;s a far clearer path to ensuring this work can be used by everyone.&lt;/p&gt;
&lt;p&gt;I&apos;ve chosen to use OxCaml both for the build environment and the runtime environment on the web. While the choice doesn&apos;t make a huge difference in the browser, a longer-term vision is to have a choice of execution environments, and oxcaml for running high-performance numerical code makes a lot of sense.&lt;/p&gt;
&lt;p&gt;The implementation work can roughly be split into three parts: &lt;a href=&quot;./#infrastructure-and-odoc&quot;&gt;Infrastructure and Odoc&lt;/a&gt;, the &lt;a href=&quot;./#x-ocaml-and-js_top_worker&quot;&gt;Js_top_worker and widgets&lt;/a&gt; libraries, and the actual &lt;a href=&quot;./#tessera-notebooks&quot;&gt;TESSERA code&lt;/a&gt;. Then I&apos;ve got some thoughts on the &lt;a href=&quot;./#implications&quot;&gt;implications&lt;/a&gt; of the experience, for open source development in general and odoc and the notebooks specifically.&lt;/p&gt;
&lt;h2&gt;Infrastructure and Odoc&lt;/h2&gt;
&lt;p&gt;In order to have a fast build cycle I needed an incremental build. Last year I wrote up my experiences writing the &lt;a href=&quot;/blog/2025/12/claude-and-dune.html&quot;&gt;dune rules for odoc&lt;/a&gt; with Claude. I opened &lt;a href=&quot;https://github.com/ocaml/dune/pull/12995&quot;&gt;the PR&lt;/a&gt;, representing a &amp;quot;feature complete&amp;quot; replacement for the current rules, in that it can completely replace what&apos;s in dune now, but doesn&apos;t extend the rules for the new features of odoc. Since then we&apos;ve merged the first part of it, and happily my colleague &lt;a href=&quot;https://choum.net/panglesd&quot;&gt;Paul-Elliot&lt;/a&gt; is back from his brief sabbatical working on &lt;a href=&quot;https://docs.slipshow.org/en/stable/&quot;&gt;slipshow&lt;/a&gt; to work on getting the rest of it merged.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;fix-doc-bugs-with-dune.gif&quot;&gt;
  &lt;figcaption&gt;&lt;em&gt;The incremental build, together with &quot;--warn-error&quot; is very helpful when fixing documentation bugs!&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;To support my goals though, I&apos;ve had to extend the rules by quite a lot from that initial work. I&apos;ve got &lt;a href=&quot;https://github.com/ocaml/dune/commit/0aa0170938b92342a609476abab67b960c356bd5&quot;&gt;support for assets&lt;/a&gt; so I can put images on my blog posts, support for &lt;a href=&quot;https://github.com/ocaml/dune/commit/91b08307c118193fc46a707891c28721c1778916&quot;&gt;source rendering&lt;/a&gt; so that you can see the code I&apos;m using. There&apos;s &lt;a href=&quot;https://github.com/ocaml/dune/commit/4cb2b33e98634eba8b9a267a6dd9bcfc959575a4&quot;&gt;markdown output&lt;/a&gt; for Claude to consume, and &lt;a href=&quot;https://github.com/ocaml/dune/commit/76f61319a21ba8b03feeb0c15e2a567ca2114489&quot;&gt;sherlodoc native support&lt;/a&gt; so you (or your agent) can run sherlodoc queries on the command line. The &lt;a href=&quot;https://github.com/ocaml/dune/commit/91b08307c118193fc46a707891c28721c1778916&quot;&gt;prefix&lt;/a&gt; for all output is configurable, so that I was able to put everything dune built under &amp;quot;/reference&amp;quot;, and pass &lt;a href=&quot;https://github.com/ocaml/dune/commit/48a5cbb79b24606304dc25f965c22aa5e4cb898e&quot;&gt;arbitrary options&lt;/a&gt; to the various invocations of odoc so that I could specify global defaults like the Javascript toplevel worker and opam repository to use in my notebooks. We&apos;ve also now got &lt;a href=&quot;https://github.com/ocaml/dune/commit/a49e98a88d7358074afffedfb0f6ee922208cdb7&quot;&gt;smarter rules&lt;/a&gt; that don&apos;t pull in as many dependencies as the current PR, so that I didn&apos;t have to install any extra packages I wasn&apos;t using just to build my own docs.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;sherlodoc-search.png&quot;&gt;
  &lt;figcaption&gt;&lt;em&gt;Sherlodoc database generation and search using dune. Note that the output is rendered with markdown!&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Most of these changes will be useful to the wider OCaml community, so it makes sense to try to upstream them. They&apos;re also fairly orthogonal to each other, and are obviously dependent on the original PR, and hence they can be made as independent PRs once the big one has been merged.&lt;/p&gt;
&lt;h3&gt;Odoc&lt;/h3&gt;
&lt;p&gt;OxCaml support for odoc was contributed by Luke Maurer early on after OxCaml was released. However, this only fixed the build of odoc, it didn&apos;t give it any mode or layouts, nor any of the other new features of OxCaml. I asked Claude to look through the way the toplevel prints these annotations and port them to odoc, and that&apos;s been implemented on this site. For example, see &lt;a href=&quot;/reference/base/base/Base/Uniform_array/index.html#val-length&quot;&gt;&lt;code&gt;Base.Uniform_array.length&lt;/code&gt;&lt;/a&gt; - you can see there the &lt;code&gt;portable&lt;/code&gt;, &lt;code&gt;local&lt;/code&gt; and &lt;code&gt;contended&lt;/code&gt; annotations on the type. If you click on the &lt;code&gt;source&lt;/code&gt; link, you&apos;ll see I also added some improvements to the source rendering - there are many more links now and we&apos;ve got the ability to link to source from doc comments and mlds.&lt;/p&gt;
&lt;figure&gt;
  &lt;a href=&quot;/reference/base/base/Base/Uniform_array/index.html#val-length&quot;&gt;&lt;img src=&quot;length-with-oxcaml.png&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;Output from odoc, showing the oxcaml extensions&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One of the earliest things I did was to give Odoc a new plugin system that has been hugely enabling for building the new features. I&apos;m using dune&apos;s &lt;a href=&quot;https://dune.readthedocs.io/en/stable/sites.html#sites&quot;&gt;site&lt;/a&gt; feature for the plugins, which really &amp;quot;just worked&amp;quot;. It was very easy both to add the feature to odoc and to create the plugins themselves. Building a whole variety of plugins has also been very useful in testing the shape of the plugin API, and I&apos;ve made numerous changes to it as I&apos;ve built them and they expose various problems.&lt;/p&gt;
&lt;p&gt;The plugins are activiated in a few different ways. The first is to allow &lt;a href=&quot;https://ocaml.org/manual/5.4/ocamldoc.html#sss:ocamldoc-custom-tags&quot;&gt;custom tags&lt;/a&gt;, e.g. &lt;code&gt;@foobar &amp;lt;content&amp;gt;&lt;/code&gt;, a feature of ocamldoc that odoc didn&apos;t previously support. The plugin registers handlers for a particular set of tags, and when odoc encounters that tag it calls the plugin to find out what should be rendered. The second way is for plugins to handle particular source-code blocks, e.g. &lt;code&gt;{@baz k1=v1 k2=v2[ ... ]}&lt;/code&gt;. Once again when odoc encounters a matching block it gets passed to the plugin to determine what should be rendered.&lt;/p&gt;
&lt;p&gt;Let&apos;s take a brief look through the plugins I&apos;ve made.&lt;/p&gt;
&lt;h4&gt;Admonitions&lt;/h4&gt;
&lt;p&gt;This is a feature we&apos;ve wanted to add to odoc for a while - and we have a &lt;a href=&quot;https://hackmd.io/ETSOAmetTI-E3vrDk3Bfrw&quot;&gt;design sketched out&lt;/a&gt; for it.&lt;/p&gt;
&lt;p&gt;NoteThis is a &amp;quot;note&amp;quot; admonition.
This is more-or-less a throwaway plugin as we&apos;ll be doing this &amp;quot;properly&amp;quot; and won&apos;t need it. It made for a nice first test though and the functionality is useful and quite important. I&apos;ve been using it to mark the truly agent-coded libraries where I&apos;ve not seen the code at all, as opposed to those where I&apos;ve been far more involved in the changes and I&apos;m slightly more confident in how they work.&lt;/p&gt;
&lt;p&gt;The code for this is &lt;a href=&quot;https://tangled.org/jon.recoil.org/odoc-admonition-extension&quot;&gt;on tangled&lt;/a&gt;, and is a good example of a very simple plugin.&lt;/p&gt;
&lt;h4&gt;Diagrams&lt;/h4&gt;
&lt;p&gt;I&apos;ve got 3 diagramming plugins - &lt;a href=&quot;/reference/odoc-mermaid-extension/&quot;&gt;odoc-mermaid-extension&lt;/a&gt;, &lt;a href=&quot;/reference/odoc-msc-extension/&quot;&gt;odoc-msc-extension&lt;/a&gt; and &lt;a href=&quot;/reference/odoc-dot-extension/&quot;&gt;odoc-dot-extension&lt;/a&gt;. These were particularly useful in understanding the need to determine the lifecycle of any javascript glue that&apos;s required, especially when we&apos;re dynamically loading pages like ocaml.org does. And in fact I&apos;ve put the Mermaid extension to use documenting the issues in the &lt;a href=&quot;/reference/nox-odoc/extensions.html&quot;&gt;extensions documentation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here&apos;s a diagram showing the dependencies of my TESSERA libraries:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;graph LR
    tessera-zarr --&amp;gt; tessera-geotessera
    tessera-zarr --&amp;gt; tessera-linalg
    tessera-zarr --&amp;gt; zarr-v3
    tessera-zarr-jsoo --&amp;gt; tessera-zarr
    tessera-zarr-jsoo --&amp;gt; zarr-v3
    tessera-geotessera --&amp;gt; tessera-linalg
    tessera-geotessera --&amp;gt; tessera-npy
    tessera-geotessera-jsoo --&amp;gt; tessera-geotessera
    tessera-tfjs --&amp;gt; tessera-linalg
    tessera-viz --&amp;gt; tessera-linalg
    tessera-viz-jsoo --&amp;gt; tessera-viz
    zarr-v3-unix --&amp;gt; zarr-v3
&lt;/code&gt;&lt;/pre&gt;
&lt;h4&gt;Interactive pages (notebooks)&lt;/h4&gt;
&lt;p&gt;My previous efforts at creating notebooks involved a separate pipeline to build them, but with this plugin it became trivial to have them built as part of the normal &apos;dune build&apos; process.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;/reference/odoc-interactive-extension/&quot;&gt;odoc-interactive-extension&lt;/a&gt; uses &lt;a href=&quot;https://github.com/art-w&quot;&gt;art-w&lt;/a&gt;&apos;s &lt;a href=&quot;https://github.com/art-w/x-ocaml&quot;&gt;x-ocaml&lt;/a&gt; to add interactivity to the mld files. The plugin itself is rather simple, really just finding marked ocaml code blocks and outputting an &lt;code&gt;&amp;lt;x-ocaml&amp;gt;&lt;/code&gt; element into the HTML. The metadata in the code block is translated directly into attributes in the x-ocaml tag allowing arbitrary parameters to be passed.&lt;/p&gt;
&lt;p&gt;I have some more detail on x-ocaml and js_top_worker &lt;a href=&quot;./#x-ocaml-and-js_top_worker&quot;&gt;later in this post&lt;/a&gt;.&lt;/p&gt;
&lt;h4&gt;Scrollycode&lt;/h4&gt;
&lt;p&gt;The &lt;a href=&quot;/reference/odoc-scrollycode-extension/&quot;&gt;odoc-scrollycode-extension&lt;/a&gt; is based on &lt;a href=&quot;https://pomb.us&quot;&gt;Rodrigo Pombo&lt;/a&gt;&apos;s work and similar scroll-linked tutorial sites like MLU&apos;s &lt;a href=&quot;https://mlu-explain.github.io/decision-tree/&quot;&gt;decision trees&lt;/a&gt; page. I initially had Claude come up with a sketch to show it working, complete with Merlin integration so that types-on-hover work:&lt;/p&gt;
&lt;figure&gt;
  &lt;a href=&quot;https://jon.ludl.am/experiments/scrollycoder/&quot;&gt;&lt;img src=&quot;claude-scrolly-sketch.gif&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;One of claude&apos;s mockups of the idea&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;With that as &amp;quot;proof-of-concept&amp;quot; I then built a new odoc plugin so that you can just insert &lt;a href=&quot;https://tangled.org/jon.recoil.org/odoc-scrollycode-extension/blob/main/doc/notebook_testing.mld#L3&quot;&gt;special markup&lt;/a&gt; in the source and you get lovely animated tutorials:&lt;/p&gt;
&lt;figure&gt;
  &lt;a href=&quot;/reference/odoc-scrollycode-extension/notebook_testing.html&quot;&gt;&lt;img src=&quot;scrolly-odoc.gif&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;The output from odoc&apos;s plugin&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The screen recording above is from one of the &lt;a href=&quot;/reference/odoc-scrollycode-extension/notebook_testing.html&quot;&gt;test pages&lt;/a&gt;, and currently shows only the incremental building of a module. Rather more interesting is something like Rodrigo&apos;s original &lt;a href=&quot;https://pomb.us/build-your-own-react/&quot;&gt;build your own react&lt;/a&gt; page that showcases &lt;em&gt;modification&lt;/em&gt; of the source as you scroll though, where it starts with a very simple program and progressively builds in more features. I&apos;d like to extend my plugin to support this soon.&lt;/p&gt;
&lt;h4&gt;HTML shells&lt;/h4&gt;
&lt;p&gt;The &amp;quot;traditional&amp;quot; way that you embed odoc output into another webpage, like ocaml.org does, is to output JSON. This works well where you&apos;ve got a smart server that can read the JSON and render it on demand, but when you&apos;re trying to build a static website like the dune rules produce, it&apos;s rather less convenient.&lt;/p&gt;
&lt;p&gt;The HTML shells extension is part of the odoc plugin system, and allows you to swap out the default HTML renderer for another completely customisable one. I have two plugins that use this system:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-docsite/&quot;&gt;odoc-docsite&lt;/a&gt; which produces a more modern SPA-style site.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-jons-plugins/&quot;&gt;jons-shell&lt;/a&gt; which produces this website.
The advantage of using these is that it doesn&apos;t affect the flow of the documentation pipeline, so there are no changes required to the dune rules for it to produce a very different output. You just install the plugin, add the &apos;--shell &amp;lt;...&amp;gt;&apos; to the dune file or dune-workspace file at the root of your repository, and run &lt;code&gt;dune build @doc&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The second of these plugins is a good example of a bit more of an involved one - see the &lt;a href=&quot;https://tangled.org/jon.recoil.org/odoc-jons-plugins&quot;&gt;source code&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;X-OCaml and Js_top_worker&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/art-w/x-ocaml&quot;&gt;x-ocaml&lt;/a&gt; is &lt;a href=&quot;https://github.com/art-w&quot;&gt;art-w&lt;/a&gt;&apos;s excellent tool to to add interactivity to any web page. It works by registering a new HTML element &lt;code&gt;x-ocaml&lt;/code&gt; that can be used like any normal HTML element, producing a Codemirror-backed editor with a &amp;quot;run&amp;quot; button that will execute your code in your browser using a web worker.&lt;/p&gt;
&lt;p&gt;In order to use OxCaml in my notebooks, I needed a stronger separation between the code running in the browser context and the &amp;quot;execution engine&amp;quot; running in the web worker. I switched from &lt;a href=&quot;/reference/ocaml-compiler/stdlib/Stdlib/Marshal/index.html&quot;&gt;&lt;code&gt;Marshal&lt;/code&gt;&lt;/a&gt; for the communications layer to &lt;a href=&quot;/reference/js_top_worker-client/js_top_worker-client.msg/Js_top_worker_client_msg/index.html&quot;&gt;JSON&lt;/a&gt;, as marshal is incompatible between OxCaml and OCaml, which meant that a single frontend implementation can now talk to both OxCaml and OCaml workers running in the web worker.&lt;/p&gt;
&lt;p&gt;My fork is using &lt;a href=&quot;https://github.com/jonludlam/js_top_worker&quot;&gt;js_top_worker&lt;/a&gt; as the library and tool that&apos;s used to manage the javascript toplevels. The goal of the tool is to use it to have toplevels or interactive notebooks for every package in opam-repostiory, so it&apos;s important that it&apos;s careful about what can be cached. As such, the design is that there&apos;s a single toplevel worker, and libraries can be dynamically loaded in via &lt;code&gt;#require&lt;/code&gt;, just like the standard toplevel. Each library is compile from a &lt;code&gt;cma&lt;/code&gt; to a &lt;code&gt;cma.js&lt;/code&gt;, and the &lt;code&gt;cmi&lt;/code&gt;s are also dynamically fetched as they&apos;re needed. The worker, the &lt;code&gt;cma.js&lt;/code&gt; files and &lt;code&gt;cmi&lt;/code&gt;s are then hosted alongside &lt;code&gt;META&lt;/code&gt; files and an index.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;/_opam/findlib_index.json&quot;&gt;Toplevel index file&lt;/a&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{&amp;quot;meta_files&amp;quot;:
  [ &amp;quot;lib/yojson/META&amp;quot;,
    &amp;quot;lib/uutf/META&amp;quot;,
    &amp;quot;lib/uri/META&amp;quot;,
    ...
    &amp;quot;lib/js_top_worker-rpc/META&amp;quot;
    ],
 &amp;quot;compiler&amp;quot;:
   {&amp;quot;worker_url&amp;quot;: &amp;quot;worker.js&amp;quot;}
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;One important new feature of js_top_worker that I&apos;ve built for the notebooks is the ability to use interactive widgets. This requires two-way communications between the main javascript thread and the web-worker backend as we want to be able to trigger changes from both sides. Rather than hard-coding all the widgets that can be made, the design allows new ones using standard OCaml libraries, so you can dynamically load in the widgets you require into your notebook, or even write new widgets there directly.&lt;/p&gt;
&lt;p&gt;Each widget either has direct DOM callbacks, or injects some javascript glue code into the frontend, which then communicates between back and frontend using JSON. This layer is mostly &amp;quot;stringly typed&amp;quot; using the &lt;a href=&quot;/reference/js_top_worker-widget/js_top_worker-widget/Widget/index.html&quot;&gt;widget library&lt;/a&gt;. On top of that you can then build more strongly typed libraries like the &lt;a href=&quot;/reference/js_top_worker-widget-leaflet/&quot;&gt;Leaflet library&lt;/a&gt; that give a more pleasant interface. Communications happen in both directions, so code can cause widgets to change, and interaction with the widgets can cause code to execute.&lt;/p&gt;
&lt;figure&gt;
  &lt;a href=&quot;/reference/js_top_worker-widget-leaflet/index.html&quot;&gt;&lt;img src=&quot;leaflet-fly.gif&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;Click to fly, executing via OCaml&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Try the demo out yourself &lt;a href=&quot;/reference/js_top_worker-widget-leaflet/&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;Execution model&lt;/h3&gt;
&lt;p&gt;My colleague Michael Dales &lt;a href=&quot;https://digitalflapjack.com/blog/marimo/&quot;&gt;wrote about&lt;/a&gt; the challenges of balancing interactivity and incremental computation in notebooks, and his post he settled on &lt;a href=&quot;https://marimo.io/&quot;&gt;Marimo&lt;/a&gt;. The approach that Marimo takes is to analyse the code and build a dependency graph so that as values are adjusted, any affected part of the computation can be re-evaluated. You can then create values with UI elements attached, and the whole experience is very pleasant.&lt;/p&gt;
&lt;p&gt;The &amp;quot;usual&amp;quot; approach to this sort of UI interaction in OCaml is to use an incremental library of some sort, like Jane Street&apos;s &lt;a href=&quot;https://github.com/janestreet/bonsai/&quot;&gt;Bonsai&lt;/a&gt; or &lt;a href=&quot;https://erratique.ch/software/react&quot;&gt;React&lt;/a&gt;. In this way we can build the dependency graph and re-evaluation semantics &lt;em&gt;within the language&lt;/em&gt;. In these examples, I&apos;ve used &lt;a href=&quot;https://erratique.ch/software/note&quot;&gt;Note&lt;/a&gt;. The interesting wrinkle in this work is that because all interactions are being mediated by &lt;code&gt;js_top_worker&lt;/code&gt;, we&apos;ve got a single place where we could &lt;em&gt;record&lt;/em&gt; the interactions that have happened, and then potentially &amp;quot;play them back&amp;quot;.&lt;/p&gt;
&lt;p&gt;With support for OxCaml, widgets and dynamic library loading, and integration into the docs pipeline via the odoc interactive plugin, it was time to create a real-world notebook to verify that it could be used to do &amp;quot;proper science&amp;quot;!&lt;/p&gt;
&lt;h2&gt;TESSERA notebooks&lt;/h2&gt;
&lt;p&gt;One of the most exciting things to come out of our group recently has been &lt;a href=&quot;https://geotessera.org/&quot;&gt;TESSERA&lt;/a&gt;, a pixel-wise Earth observation foundation model. The code and demos for this are mostly in Python, particularly using Jupyter notebooks. This provided me with an excellent stress test that tied together many of the strands of this work and validated that they could be used together to build something useful. A &lt;a href=&quot;https://github.com/ucam-eo/tessera-interactive-map&quot;&gt;simple python notebook&lt;/a&gt; showing the basic workflow already existed, so I decided to port it to OxCaml, and more specifically to make an odoc notebook. You can run the resulting &lt;a href=&quot;/notebooks/interactive_map.html&quot;&gt;OxCaml notebook&lt;/a&gt; in your own browser, with no server required at all.&lt;/p&gt;
&lt;p&gt;The notebook works by first presenting a map, and asking the user to select an area to download the TESSERA embeddings. Once we&apos;ve got the 128-dimensional per-pixel values, we then need to display them. Rather than simply picking 3 dimensions as our RGB values, we do a PCA computation to choose some vectors for our colours. This is then overlaid on the map by first turning it into a &lt;code&gt;.png&lt;/code&gt;, then encoding it as a base64 data URL, and passing it to the frontend to overlay on the map. The user then picks some points on the map, categorising them into arbitrary groups. We then run a knn classifier over the whole downloaded area and finally overlay that image.&lt;/p&gt;
&lt;p&gt;I struggled a little with co-ordinate systems, mostly because the distortion if you &lt;em&gt;dont&lt;/em&gt; map is fairly small, so it looks pretty good. However, you can see clearly on this image that the A14 (the red road at the bottom left) is a little bit off:&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;tessera.png&quot;&gt;
  &lt;figcaption&gt;&lt;em&gt;An earlier attempt, showing the distortion&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In the newer notebook where I do the remapping correctly, it matches much better. You can see the remapping by looking at the left and right sides of the overlay, where they&apos;re no longer quite vertical.&lt;/p&gt;
&lt;figure&gt;
  &lt;a href=&quot;/notebooks/interactive_map.html&quot;&gt;&lt;img src=&quot;notebook.png&quot; alt=&quot;A client-side notebook using TESSERA&quot;&gt;&lt;/a&gt;
  &lt;figcaption&gt;&lt;em&gt;Remapped overlay showing better matching with the map&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This version uses the &lt;a href=&quot;/reference/tessera-npy/tessera-npy/Npy/index.html&quot;&gt;Numpy&lt;/a&gt; interface to TESSERA, which downloads the embeddings in quite large (100M) chunks. While it was nice that this worked, it meant it was quite slow to see anything actually happening. I then switched to using the &lt;a href=&quot;/reference/zarr-v3/zarr-v3/Zarr_v3/index.html&quot;&gt;&lt;code&gt;Zarr_v3&lt;/code&gt;&lt;/a&gt; format, which is far more efficient, using range requests, compression and better encoding to reduce the bandwidth requirements. This version is &lt;a href=&quot;/notebooks/interactive_map_zarr.html&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This really demonstrated that the approaches I&apos;ve made with the various different strands of work can all be knitted together in a very useful way, allowing us to bring the strong type system of OCaml together with the ubiquitous runtime of the browser, and using the power of WebGPU to run calculations that can help to change our world for the better.&lt;/p&gt;
&lt;h2&gt;Implications&lt;/h2&gt;
&lt;p&gt;I&apos;ve written (or caused to be written) a lot of code for this project. It&apos;s all &amp;quot;open source&amp;quot; in that the code is available, though this is the very minimalist sense of the term. I&apos;ve both modified other people&apos;s code, and written other things from scratch, and pushed everything publicly on the &lt;a href=&quot;https://tangled.org/jon.recoil.org/monopam-myspace&quot;&gt;repostiory&lt;/a&gt;. It&apos;s been a fundamentally different experience than everything I&apos;ve done in the past 25 years or so of OSS development.&lt;/p&gt;
&lt;p&gt;Traditionally with open source software, most people never change the code. The vast majority of people take the code, run it, use it, and get on with their lives, grateful to varying degrees to the authors. A much smaller number of people will make changes to the code. This requires an enormous investment of effort, firstly in learning to code at all and then learning the specific codebase sufficiently well to make a change. A small subset of these people go to the additional effort of trying to upstream the code, and open a PR. At this point, the maintainers of the code now &lt;em&gt;also&lt;/em&gt; need to put in a lot of effort to review what&apos;s been sent, and engage with the contributor to address any issues with the code. So it&apos;s hard work. Why do people bother?&lt;/p&gt;
&lt;p&gt;This process is helpful to the contributor as maintaining patches on an open source project is not only work, but &lt;em&gt;unpredictable&lt;/em&gt; work. Upstream changes can require you to have to fix your modifications at any time, and those fixes might be anywhere from trivial to completely impossible. Getting your patch upstreamed means that there&apos;s a much better chance that whatever it is that was fixed or improved remains that way without further effort.&lt;/p&gt;
&lt;p&gt;On the part of the maintainer, the effort of reviewing and interacting with the contributor will hopefully both improve the code, but also effectively enlarge your workforce as the contributor learns more about the project. The more people who know about the project, the easier maintenance becomes.&lt;/p&gt;
&lt;p&gt;LLMs have totally upset the balance of these motivations. From the contributor&apos;s perspective, the effort now required to maintain your own fork of a project is &lt;em&gt;much&lt;/em&gt; smaller, as you don&apos;t need to learn the codebase anywhere near as deeply, and you don&apos;t have to worry anywhere near as much about maintaining your fork as upstream changes. So it&apos;s both easier for there to be code changes, and it&apos;s less rewarding to try and upstream them. From the perspective of the maintainer, there may be many more low quality pull requests, and in many cases there will be limited value in trying to cultivate new contributors as the LLMs can&apos;t learn directly from the feedback, and the humans guiding the LLMs will know much less about the project as they haven&apos;t needed to in order to make the PR.&lt;/p&gt;
&lt;p&gt;Where these adjusted values will lead is very hard to predict. Empirically though, I find that I&apos;ve made few PRs over the last few months, but I&apos;ve amassed a lot of new repositories and forks of existing repos, and it&apos;s been very easy to slip into this mode of working. The effect on non-programmers hasn&apos;t begun to happen really, but there &lt;em&gt;will&lt;/em&gt; be an effect and it &lt;em&gt;will&lt;/em&gt; be profound. I&apos;ll know when that moment arrives when my artist-rather-than-scientist family members are using LLMs to alter code - which to them will just be the latest change in MacOS that&apos;s happened, but &lt;em&gt;I&apos;ll&lt;/em&gt; know what&apos;s going on.&lt;/p&gt;
&lt;h2&gt;Attribution&lt;/h2&gt;
&lt;p&gt;My early purely-agentic commits were all authored by me and co-authored by Claude. This is a lie. Morally, I found I couldn&apos;t carry on like this so I&apos;ve switched now to having the commits authored by &amp;quot;Jon&apos;s Agent&amp;quot;. My plan is to rewrite the author when I&apos;ve gone through it line by line, and even then, I feel that Claude ought to be marked as Author and I should just be adding my &amp;quot;Reviewed-by&amp;quot; line onto it. In either case, for a pull request to be made the human in the loop has the responsibility to justify the changes, and therefore has to be totally familiar with both the changes and the code being changed. I&apos;m not at all fundamentally opposed to having LLMs involved in the process of making changes to open source code, but in my experience so far, there&apos;s a huge amount of effort that needs to go into it even after you&apos;ve got working code and tests. I&apos;ve tested the waters with my &lt;a href=&quot;https://github.com/ocaml/dune/pull/12995&quot;&gt;dune PR&lt;/a&gt;, and have merged a &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1402&quot;&gt;simple bugfix or two&lt;/a&gt; to odoc. Even these one-liners needed careful thought and attention before I felt I could make a PR though, and the dune PR took &lt;a href=&quot;/blog/2025/12/claude-and-dune.html&quot;&gt;much more effort than that&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Bug-discovery&lt;/h2&gt;
&lt;p&gt;One thing that I&apos;ve found tremendously useful is narrowing down bugs. Armed with a repro, setting Claude off to track down issues has been a wonderful time saver. Additionally, asking it to explain the issue in detail with links to the source is very handy indeed. It&apos;s not, however, able to discover &lt;em&gt;all&lt;/em&gt; bugs, even with a lot of time. When working on the fix for a &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1400&quot;&gt;particularly nasty bug&lt;/a&gt;, I found that with the patch applied we&apos;d get a different error somewhere deep in some of Jane Street&apos;s async ecosystem. I had a good suspicion of what the problem was given the changes that had been made already, as the code I had altered had an analogue elsewhere in the codebase that hadn&apos;t been fixed, so I thought this would be quite a good test for Claude. I gave it lots of hints, but it flailed at the problem for several hours, often giving up, sometimes blaming the OxCaml compiler, and sometimes upstream OCaml. In the end I couldn&apos;t get Claude to make useful progress so I implemented what I thought &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1400/changes/3164566c010cd110b283643cdf18cbb9ab3400c4&quot;&gt;should be the fix&lt;/a&gt; and indeed the problem went away. So there are clearly classes of bugs that are still beyond Claude&apos;s capabilities to understand, which isn&apos;t an unexpected result.&lt;/p&gt;
&lt;h2&gt;API boundaries&lt;/h2&gt;
&lt;p&gt;I&apos;ve found it very helpful to have API boundaries to help structure the code that Claude has been producing. Anil has long been enthusiastically pushing the idea that we should write the &lt;code&gt;mli&lt;/code&gt; files first, which constrain what types and values are available between the modules. We can then write tests that target these interfaces, and then adjust them where the tests have shown them to be inelegant or downright unusable. We can then write the implementations and watch the tests start to pass. A particularly interesting example of this is the odoc plugin interface. The experience of writing several very different plugins that all extended odoc in different directions was very helpful, and I adjusted the interface quite a few times. I also adjusted the &lt;a href=&quot;/reference/nox-odoc/extensions.html&quot;&gt;documentation&lt;/a&gt;, where qualities about how the interface &lt;em&gt;behaved&lt;/em&gt; that weren&apos;t obvious from the types could be carefully noted, for example how scripts might be made to behave correctly when the odoc pages were in an SPA shell. Far too often in the OCaml community we&apos;ve relied just on types and modules as documentation for our libraries, and this is rarely enough.&lt;/p&gt;
&lt;h2&gt;Failure modes&lt;/h2&gt;
&lt;p&gt;When I was working on the &lt;a href=&quot;/blog/2025/12/claude-and-dune.html&quot;&gt;dune rules&lt;/a&gt;, I made the mistake of going too long without giving Claude some architectural constraints, and I ended up with a Big Ball of Code that I then had to spend time unpicking and teasing apart into sensible looking modules. I had rather hopefully believed this to be a Claude Opus pre-4.5 problem, but I still hit this recently when adding a new feature, when it just added vast amounts of code to one file to implement it, and the code was all very unstructured and unsatisfactory. This was despite going through a design process where we went through the goals, the use cases and desired features, but crucially, not at the level of the code.&lt;/p&gt;
&lt;p&gt;Another failure mode I observed was &lt;em&gt;my&lt;/em&gt; failure. It&apos;s very easy, and very tempting, to get your agent to do the next neat thing on the roadmap. Especially when you&apos;ve just spent a while going through the design and planning for the previous feature and Claude has got started on it. The problem is that this can generate a large amount of code that kind-of-works but has a bunch of bugs, which can end up costing a lot more time and effort. The cost of starting the agent going is much smaller than the cost of wading through the results, and it&apos;s quite easy to end up drowning under a load of very interesting and partly cool half results. I&apos;m very much reminded of Dr Ian Malcolm&apos;s words from Jurassic Park: &amp;quot;your scientists were so preoccupied with whether or not they could, they didn&apos;t stop to think if they should.&amp;quot;&lt;/p&gt;
&lt;h2&gt;What&apos;s next?&lt;/h2&gt;
&lt;p&gt;In the more immediate future though, perhaps with my odoc-maintainer hat on, I still want to make sure that I extract the useful parts of what I&apos;ve done and get them upstreamed. While everything I&apos;ve done is all open source and published, there&apos;s a lot of work I need to do so that others in the community will benefit from my changes. While it&apos;s technically possible to add my opam-repo to your switch, and install my versions of odoc, dune, and my various plugins, I don&apos;t think many people are actually going to do that. Worse than that, people might get their agents to just grab the source and mutate it further, just diluting the efforts going into it.&lt;/p&gt;
&lt;p&gt;Fortunately I&apos;ve been talking with &lt;a href=&quot;https://choum.net/panglesd/&quot;&gt;Paul Elliot&lt;/a&gt;, who has volunteered to shepherd the dune PR through to completion. I&apos;ll be working with him on this of course, but I&apos;m hoping he&apos;ll be doing the lion&apos;s share of the work.&lt;/p&gt;
&lt;p&gt;The OxCaml work will be taken on by &lt;a href=&quot;https://github.com/art-w&quot;&gt;art-w&lt;/a&gt; who&apos;s already done an excellent job getting Luke Maurer&apos;s patches into shape and PR&apos;d to ocaml/odoc.&lt;/p&gt;
&lt;p&gt;We&apos;ll be having a forward-looking odoc meeting on the 8th April to think about the roadmap for odoc, and I&apos;m confident that my work will help us in the discussions of what directions odoc will be heading in over the next year or so. Whilst the bigger picture of how the world is evolving plays out we need to at least carry on maintaining the community software, and use LLMs wisely and carefully as another tool in our toolbox to help us keep the code as useful and bug free as possible.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/04/odoc_and_ocaml_notebooks.html" rel="alternate" title="Odoc and OCaml Notebooks"/>
    <category term="odoc"/>
    <category term="notebooks"/>
    <category term="tessera"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/03/weeknotes-2026-13.html</id>
    <title type="text">Weeknotes 2026 week 13</title>
    <updated>2026-03-31T00:00:00Z</updated>
    <published>2026-03-31T00:00:00Z</published>
    <content type="html">
      &lt;h2&gt;What did I do?&lt;/h2&gt;
&lt;p&gt;I spent rather a long time this week working an a review of the past few months of work, writing, rewriting, checking what I&apos;ve written, rearranging, scratching my head and pondering. Too long. I did do some other stuff too though:&lt;/p&gt;
&lt;h3&gt;Standalone TESSERA page&lt;/h3&gt;
&lt;p&gt;I had built the Js_top_worker code with the intention of having it runnable in any old web page, not necessarily one hosted on the same site as the code. However, I hadn&apos;t actually demonstrated this working, so I set up a tangled-hosted site at &lt;a href=&quot;https://jonludlam.tngl.io&quot;&gt;https://jonludlam.tngl.io&lt;/a&gt; and put a basic Tessera notebook there. This (fortunately) worked more-or-less first time.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;tessera-standalone.png&quot;&gt;
  &lt;figcaption&gt;&lt;em&gt;Standalone TESSERA notebook&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The source for this is just a single html page. The repository is on &lt;a href=&quot;https://tangled.org/jon.recoil.org/jonludlam.tngl.io/&quot;&gt;tangled&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;Oxmono docs build&lt;/h3&gt;
&lt;p&gt;Anil pointed out that his docs build in &lt;a href=&quot;https://github.com/avsm/oxmono&quot;&gt;oxmono&lt;/a&gt; had stopped working - or perhaps never started working? I&apos;d looked at this a couple of weeks back and found it was to do with the source rendering, which wasn&apos;t working correctly in the presence of multiple implementations of virtual libraries. I had &amp;quot;fixed&amp;quot; that on the 3.22 branch by disabling the source rendering rules by default, but I found out subsequently that dune 3.22 had some issues with oxcaml, so I needed to get the fix back onto the 3.21.1 branch. With that done, I tried running the docs build again, but it&apos;s OOMing on the rather limited (7G) workers that we get for free with GHA. Now it&apos;s definitely suspect that we can build the code itself but not the docs, so there&apos;s some odoc work to do there, but we should probably just build it on one of our servers before we tackle that.&lt;/p&gt;
&lt;h3&gt;Potential scrolly improvements&lt;/h3&gt;
&lt;p&gt;While writing the review, I had another look at the examples I&apos;ve got in the &lt;a href=&quot;/reference/odoc-scrollycode-extension/&quot;&gt;Scrollycode plugin&lt;/a&gt; and rediscovered that they are quite underwhelming. They&apos;re obviously Claude-generated placeholders, but I&apos;ve always had a particular use-case in mind. About 10 years ago I gave a series of seminars at Citrix, where, using the database code as example, I showed how we can use GADTs to improve upon what we had. The talks were structured by starting with some very simple code and improving it gradually, in much the same way that the &lt;a href=&quot;https://pomb.us/build-your-own-react/&quot;&gt;build-your-own-react&lt;/a&gt; example of scrollycode does. This seems to me a good way to validate the odoc plugin, so I&apos;ve dug that out. What&apos;s missing from the plugin, though, is &lt;em&gt;modification&lt;/em&gt; support - currently we can only &lt;em&gt;add&lt;/em&gt; code, so that will need to be fixed first.&lt;/p&gt;
&lt;h2&gt;What&apos;s next?&lt;/h2&gt;
&lt;p&gt;I&apos;ve more-or-less achieve what I set out to do with the TESSERA notebook - I&apos;ve demonstrated that OCaml is a perfectly capable platform for using in a notebook context to do purely client-side processing of geospatial data. For the next steps, I&apos;m going to need to talk to my colleagues to figure out what the most useful thing to do is for TESSERA. There&apos;s plenty to do, but to prioritise everything correctly is going to need some discussion.&lt;/p&gt;
&lt;p&gt;There are some open questions / areas to investigate in the purely notebook space:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The worker and cma.js/cmi files have a dependency tree. Maybe it&apos;d be interesting to publish the metadata for these using the at:// protocol?&lt;/li&gt;
&lt;li&gt;Better discoverability. You can &apos;#require&apos; things, but what&apos;s available?&lt;/li&gt;
&lt;li&gt;Docs links. Each of the packages I&apos;ve published has the docs available. It&apos;d be nice to link them up somehow.&lt;/li&gt;
&lt;li&gt;AI agent integration. Most of the notebooks I&apos;ve &amp;quot;written&amp;quot; recently have been generated by Claude. Can you paste in a token and get claude in your notebooks?&lt;/li&gt;
&lt;li&gt;Reuse/remix of notebooks. The &lt;a href=&quot;https://dl.acm.org/doi/10.1145/3759536.3763802&quot;&gt;Fairground&lt;/a&gt; paper emphasises that you can treat your notebooks as &lt;em&gt;libraries&lt;/em&gt;. What does this look like for these?&lt;/li&gt;
&lt;li&gt;Persistence? The exam-question notebooks in the examples do persist your edits right now, but that&apos;s ad-hoc and specific. What&apos;s the longer-term story?&lt;/li&gt;
&lt;li&gt;Editing beyond the code blocks. What about editing the prose in the browser?
Some more thoughts. Marimo uses python source as the primary source of the notebooks. One easy thing to do would be to have an &lt;em&gt;ml&lt;/em&gt; file as the primary source for these. We&apos;ve got easy access to the comments, which are already written in odoc markup. If we did this we&apos;d get the power of merlin in the editor, which would be pretty nice.&lt;/li&gt;
&lt;/ul&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/03/weeknotes-2026-13.html" rel="alternate" title="Weeknotes 2026 week 13"/>
    <category term="weeknotes"/>
    <category term="tessera"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/03/weeknotes-2026-12.html</id>
    <title type="text">Weeknotes 2026 week 12</title>
    <updated>2026-03-23T00:00:00Z</updated>
    <published>2026-03-23T00:00:00Z</published>
    <content type="html">
      &lt;h2&gt;What did I do?&lt;/h2&gt;
&lt;p&gt;End of term this week, so my tutorial interviews kept me quite busy for some of the week. Then Paul-Elliot has returned from his sojourn working on Slipshow, so I had a useful few hours working with him to try and begin the process of handing over the dune odoc PR.&lt;/p&gt;
&lt;p&gt;I mentioned &lt;a href=&quot;/blog/2026/03/weeknotes-2026-11.html&quot;&gt;last week&lt;/a&gt; that the TESSERA reprojection was causing issues with the overlay alignment. Now the original loading of the patches was taking ages, so I switched to zarr to be more efficient. Also, the PCA was slow, so I switched that over to tensorflow.js to use the GPU.&lt;/p&gt;
&lt;p&gt;With these two optimisations in place, it was now a lot quicker to see if the misalignment was still there. It seemed likely that the issue was translating between the UTM grid that TESSERA uses and the WGS84 coordinate system that Leaflet.js is using. It was quite quick to whip up a &lt;a href=&quot;/reference/tessera-geotessera/tessera-geotessera/Geotessera/Utm/index.html#val-wgs84_to_utm&quot;&gt;conversion routine&lt;/a&gt; (click on &apos;source&apos; to see the gory details!). With that in place, the overlay now matches up precisely as you can see on the map below. Be patient while it runs, you&apos;ll see the overlay on the map in a few seconds!&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;#require &amp;quot;tessera-zarr-jsoo&amp;quot;;;
#require &amp;quot;tessera-viz-jsoo&amp;quot;;;
#require &amp;quot;tessera-tfjs&amp;quot;;;
#require &amp;quot;js_top_worker-widget-leaflet&amp;quot;;;
open Widget_leaflet;;
register ();;
(* Load fzstd (Zstd decompressor) and TensorFlow.js *)
let () =
let open Js_of_ocaml in
let import url : unit = Js.Unsafe.fun_call
(Js.Unsafe.get Js.Unsafe.global (Js.string &amp;quot;importScripts&amp;quot;))
[| Js.Unsafe.inject (Js.string url) |] in
import &amp;quot;https://cdn.jsdelivr.net/npm/fzstd@0.1.1/umd/index.js&amp;quot;;
import &amp;quot;https://cdn.jsdelivr.net/npm/@tensorflow/tfjs@4/dist/tf.min.js&amp;quot;&lt;/x-ocaml&gt;&lt;/p&gt;
&lt;h3&gt;Reprojection test&lt;/h3&gt;
&lt;p&gt;Create the map, fetch embeddings from the Zarr store, run PCA via TensorFlow.js, and overlay — all in one cell so the async pipeline completes before the overlay is drawn:&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot; run-on=&quot;load&quot;&gt;let status_view text =
let open Widget.View in
Element { tag = &amp;quot;div&amp;quot;; attrs = [Style (&amp;quot;padding&amp;quot;, &amp;quot;8px&amp;quot;); Style (&amp;quot;font-family&amp;quot;, &amp;quot;monospace&amp;quot;)];
children = [Text text] }&lt;/p&gt;
&lt;p&gt;let () = Widget.display ~id:&amp;quot;status&amp;quot; ~handlers:[] (status_view &amp;quot;Initialising...&amp;quot;)&lt;/p&gt;
&lt;p&gt;let map = Leaflet_map.create
~center:(52.30690, -0.03296) ~zoom:14 ~height:&amp;quot;500px&amp;quot;
()&lt;/p&gt;
&lt;p&gt;let bbox = Geotessera.{
min_lat = 52.29924; min_lon = -0.05845;
max_lat = 52.31745; max_lon = -0.00755;
}&lt;/p&gt;
&lt;p&gt;let downsample mat ~h ~w ~max_pixels =
let n = h * w in
if n &amp;lt;= max_pixels then (mat, h, w)
else
let stride = int_of_float (ceil (sqrt (float_of_int n /. float_of_int max_pixels))) in
let h&apos; = (h + stride - 1) / stride in
let w&apos; = (w + stride - 1) / stride in
let out = Linalg.create_mat ~rows:(h&apos; * w&apos;) ~cols:mat.Linalg.cols in
for i = 0 to h&apos; - 1 do
for j = 0 to w&apos; - 1 do
let si = min (i * stride) (h - 1) in
let sj = min (j * stride) (w - 1) in
for f = 0 to mat.Linalg.cols - 1 do
Linalg.mat_set out (i * w&apos; + j) f
(Linalg.mat_get mat (si * w + sj) f)
done
done
done;
(out, h&apos;, w&apos;)&lt;/p&gt;
&lt;p&gt;let () =
Widget.update ~id:&amp;quot;status&amp;quot; (status_view &amp;quot;Opening Zarr store...&amp;quot;);
Lwt.async (fun () -&amp;gt;
let open Lwt.Syntax in
let* store = Tessera_zarr_jsoo.open_store () in
let progress msg = Widget.update ~id:&amp;quot;status&amp;quot; (status_view msg) in
let* (mat_full, h_full, w_full, geo_bounds) =
Tessera_zarr.fetch_region ~progress ~store bbox in
Widget.update ~id:&amp;quot;status&amp;quot;
(status_view (Printf.sprintf &amp;quot;Fetched %d×%d. Downsampling...&amp;quot; h_full w_full));
let (mat, h, w) = downsample mat_full ~h:h_full ~w:w_full ~max_pixels:500_000 in
let bounds = Leaflet_map.{
south = geo_bounds.Geotessera.min_lat;
north = geo_bounds.Geotessera.max_lat;
west = geo_bounds.Geotessera.min_lon;
east = geo_bounds.Geotessera.max_lon;
} in
Widget.update ~id:&amp;quot;status&amp;quot;
(status_view (Printf.sprintf &amp;quot;Computing PCA on %d×%d mosaic...&amp;quot; h w));
let proj = Tfjs.pca mat ~n_components:3 in
let img = Viz.pca_to_rgba ~width:w ~height:h proj in
let url = Viz_jsoo.to_data_url img in
Leaflet_map.add_image_overlay map ~url ~bounds ~opacity:0.7 ();
Widget.update ~id:&amp;quot;status&amp;quot;
(status_view (Printf.sprintf
&amp;quot;Done. Input bbox: S%.5f W%.5f N%.5f E%.5f | Overlay bounds: S%.5f W%.5f N%.5f E%.5f&amp;quot;
bbox.min_lat bbox.min_lon bbox.max_lat bbox.max_lon
bounds.south bounds.west bounds.north bounds.east));
Lwt.return_unit)&lt;/x-ocaml&gt;&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/03/weeknotes-2026-12.html" rel="alternate" title="Weeknotes 2026 week 12"/>
    <category term="weeknotes"/>
    <category term="tessera"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/03/weeknotes-2026-11.html</id>
    <title type="text">Weeknotes 2026 week 11</title>
    <updated>2026-03-18T00:00:00Z</updated>
    <published>2026-03-18T00:00:00Z</published>
    <content type="html">
      &lt;h2&gt;What did I do?&lt;/h2&gt;
&lt;h3&gt;TESSERA&lt;/h3&gt;
&lt;p&gt;I Started looking at fixing my TESSERA notebook to make it actually correct:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Removing the OCaml native PCA to replace with a tensorflow.js one. This should be much faster and probably more accurate.&lt;/li&gt;
&lt;li&gt;Investigating the slightly-off-seeming embeddings - you can see features like roads on the embeddings and they don&apos;t &lt;em&gt;quite&lt;/em&gt; match up with the map. The problem is to do with WGS84 vs UTM, which I need sort out.&lt;/li&gt;
&lt;li&gt;Looking at UNets as part of replicating Sadiq&apos;s &lt;a href=&quot;https://github.com/sadiqj/tessera-cnn-example&quot;&gt;Solar Panel&lt;/a&gt; detection example.
Not much to show yet though.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Odoc release&lt;/h3&gt;
&lt;p&gt;We need an Odoc release soon, partly to support markdown output in dune, and partly to fix some bugs that were discovered when doing the odoc v3 dune rules. I merged a lot of stuff this week, but we&apos;ve still got Arthur&apos;s OxCaml patches to merge, which are waiting on a test fix. I also looked into OCaml 5.5 compatibility. The big item looks to be the merge of Modular Explicits, which will require a bit of work to support. I&apos;m considering a mechanical fix to at least get odoc compiling with 5.5 for the odoc release, and a larger release later to properly support the 5.5 features.&lt;/p&gt;
&lt;h3&gt;OxCaml docs builds&lt;/h3&gt;
&lt;p&gt;Anil has merged &lt;a href=&quot;https://github.com/avsm/oxmono/pull/3&quot;&gt;my PR&lt;/a&gt; to his oxmono repo to do docs, and that works nicely if you&apos;re generating &lt;em&gt;all&lt;/em&gt; of the docs for your monorepo, including dependencies, which is what you get when you run &lt;code&gt;dune build @doc-full&lt;/code&gt;. Running &lt;code&gt;dune build @doc&lt;/code&gt; doesn&apos;t quite work as there&apos;s no external place like ocaml.org where you can expect your dependency libraries docs to be. Mark&apos;s &lt;code&gt;day10&lt;/code&gt; project is already building all of the oxcaml packages, so I had a look to see what it would be like to extend that to build the docs too. This is more-or-less working, but not yet as a reliable place that you can expect to be up to date. I had a thought that keeping the docs build running nicely is something that an AI agent should be able to do really well, so I&apos;m investigating this. A neat thing it&apos;s already been able to do is to sort the build problems by number of other packages that they block, and numbers one and two were PPX issues (ppx_deriving_yojson and lwt_ppx). It was then able to come up with some patches that enabled both of these to build, unblocking all of those downstream packages. This, of course, is of wider interest than just the docs, so I need to figure out what to do with those patches.&lt;/p&gt;
&lt;h3&gt;Bugfixing my monorepo&lt;/h3&gt;
&lt;p&gt;I&apos;ve been polishing the monorepo that&apos;s powering my website with a view to publishing a retrospective of the work that I&apos;ve been doing over the past few months. Most importantly, adding &apos;--warn-error&apos; to all my odoc invocations to actually notice all the warnings it&apos;s been diligently reporting, and I&apos;ve been ignoring. Looking further than my own failings though, this has really helped solidify in my mind the areas where Claude excels and the things we should be much more wary of letting it do. More on this in the retrospective!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/03/weeknotes-2026-11.html" rel="alternate" title="Weeknotes 2026 week 11"/>
    <category term="weeknotes"/>
    <category term="tessera"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/03/weeknotes-2026-10.html</id>
    <title type="text">Weeknotes 2026 week 10</title>
    <updated>2026-03-09T00:00:00Z</updated>
    <published>2026-03-09T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Here are my weeknotes for the last week, while I&apos;m still writing up some more focused posts on some specific topics - like the experience of putting everything in a monorepo to create this site, and more notes on Claude and Agentic coding in general, and its impact on the world of software. But for now, here&apos;s what I&apos;ve been up to.&lt;/p&gt;
&lt;h2&gt;What did I do?&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;New site design. The old site was a bit of a mess and was simply reusing odoc&apos;s default default styling. I&apos;ve also rearranged the content a bit to make it more navigable and cohesive.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./old.png&quot; alt=&quot;old.png&quot; &gt;
&lt;img src=&quot;./new.png&quot; alt=&quot;new.png&quot; &gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;TESSERA in the browser is a &lt;a href=&quot;https://tee.cl.cam.ac.uk/&quot;&gt;hot&lt;/a&gt; &lt;a href=&quot;https://anil.recoil.org/notes/2026w10&quot;&gt;topic&lt;/a&gt; right now, so I&apos;ve applied the work I&apos;ve been doing with x-ocaml, js_top_worker and odoc plugins to make a &lt;a href=&quot;/notebooks/interactive_map.html&quot;&gt;TESSERA notebook&lt;/a&gt; that&apos;s based on the &lt;a href=&quot;https://github.com/ucam-eo/tessera-interactive-map&quot;&gt;example notebook&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./tessera.png&quot; alt=&quot;tessera.png&quot; &gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;I was interested in whether we&apos;ll be able to do inference in reasonable time using these notebooks. &lt;a href=&quot;https://onnx.ai/&quot;&gt;ONNX&lt;/a&gt; has a web version of its runtime, so I got Claude to make some bindings, and checked it was working by doing a sentiment analysis notebook. This is working nicely, so the next step is to do something a bit more useful. Try it &lt;a href=&quot;/reference/onnxrt/sentiment_example.html&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The docs CI was again causing problems. This time it had decided that it had never built anything, and therefore needed to rebuilt the entire world. However, despite being set up as a custom dedicated runner, all its jobs were queued waiting to start. It turned out that the runner paused itself when the docker partition reached 70%. This was a little surprising on two counts - firstly we don&apos;t actually use docker for running the jobs, we use obuilder, which doesn&apos;t share space with docker. Secondly, with that in mind, how did it get to 70%? It turned out to be the job logs - including 250 gigs of older logs from a previous instance. Simply blowing those away caused everything to restart and so it&apos;s now live again.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;I met up with &lt;a href=&quot;https://ancazugo.github.io/&quot;&gt;Andrés C. Zúñiga-González&lt;/a&gt; to have a chat about how he&apos;s using interactive maps and notebooks. He pointed me at his &lt;a href=&quot;https://ancazugo.github.io/blog.html&quot;&gt;blog&lt;/a&gt;, some of which which is using &lt;a href=&quot;https://quarto.org/&quot;&gt;quarto&lt;/a&gt;, which he rates very highly. An &lt;a href=&quot;https://ancazugo.github.io/posts/2025-11-16-tessera_example.html&quot;&gt;example of quarto output&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Our group seminar this week was &lt;a href=&quot;https://tombearpark.com/&quot;&gt;Tom Bearpark&lt;/a&gt; who talked about his proposed &apos;Carbon at Risk&apos; measure in order to compare diverse ways of removing carbon from the atomsphere to help with the carbon removal market.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;What&apos;s next?&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;More writing before more coding, I think.&lt;/li&gt;
&lt;/ul&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/03/weeknotes-2026-10.html" rel="alternate" title="Weeknotes 2026 week 10"/>
    <category term="weeknotes"/>
    <category term="tessera"/>
    <category term="notebooks"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/03/weeknotes-2026-09.html</id>
    <title type="text">Weeknotes 2026 week 9</title>
    <updated>2026-03-02T00:00:00Z</updated>
    <published>2026-03-02T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Let&apos;s make this really terse!&lt;/p&gt;
&lt;h2&gt;What did I do?&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Got docs working with github actions on Anil&apos;s oxmono monorepo. Results are &lt;a href=&quot;https://jonludlam.github.io/oxmono/&quot;&gt;here&lt;/a&gt;. This includes experimental support for oxcaml modes/layouts.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Got markdown mode output into Sherlodoc&apos;s db so you can query it - great for agents!&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./search.png&quot; alt=&quot;search.png&quot; &gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Widgets in the JS OCaml toplevels - using FRP for the interactions. The neat thing here is that using FRP via Daniel Bunzli&apos;s &lt;a href=&quot;https://erratique.ch/software/note&quot;&gt;note&lt;/a&gt; library is that all the interactions are all purely functional, no refs or mutables in sight. You provide a little wrapper scripts that&apos;s run in the frontend and the interactions and send back and forth with the worker running the code where it&apos;s translated into Events and Signals. My proof-of-concept of this is a widget that works with the &lt;a href=&quot;https://leafletjs.com/&quot;&gt;leaflet.js&lt;/a&gt; library:&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;video controls src=&quot;./mapdemo.mov&quot;&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Demo coming soon!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Consolidating all of the Odoc toplevel bits and pieces into the one monorepo. Again, demo of this coming soon!&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;What am I going to do?&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;New website!&lt;/li&gt;
&lt;li&gt;Odoc plugins showcase&lt;/li&gt;
&lt;li&gt;Writing writing writing writing&lt;/li&gt;
&lt;/ul&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/03/weeknotes-2026-09.html" rel="alternate" title="Weeknotes 2026 week 9"/>
    <category term="weeknotes"/>
    <category term="odoc"/>
    <category term="plugins"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/02/weeknotes-2026-08.html</id>
    <title type="text">Weeknotes weeks 7-8</title>
    <updated>2026-02-24T00:00:00Z</updated>
    <published>2026-02-24T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;A combination one again as I took some time off due to school half term.&lt;/p&gt;
&lt;h2&gt;Finished off my exam questions&lt;/h2&gt;
&lt;p&gt;This was a lot of fun! Obviously I can&apos;t talk about it, but while it was stressful and worrying and anxiety inducing and scary, it was also engaging and interesting and thought-provoking. Having some ideas come together to make a nice coherent whole was very cool.&lt;/p&gt;
&lt;h2&gt;Testing LLMs on past paper questions&lt;/h2&gt;
&lt;p&gt;Similar to our work on the ticks that &lt;a href=&quot;https://www.youtube.com/watch?v=Ub8k1BcSRLQ&quot;&gt;Sadiq, I and others did last year&lt;/a&gt;, I wanted to try to see how well LLMs could answer tripos questions. Partly I wanted to do this so I could check that my own questions were of the right sort of level, and partly it was just a displacement activity while I wasn&apos;t making progress on the actual exam questions! I&apos;ve not done a useful analysis of the results yet, but seemed in line with our experience with the ticks, though the pass rate was lower for the same models (qwen).&lt;/p&gt;
&lt;h2&gt;Claude from a sunbed&lt;/h2&gt;
&lt;p&gt;I went away for a vitamin-D boosting bit of sun. Before I went, I got Claude to spin me up a little Telegram bridge so that I could tell it what to do, while it&apos;s still running in safeties-off mode on my sacrificial VM. This was kind of fun - I got to just indulge thoughts as they came to me, and off it would go and do stuff. It was a bit limited in how it talked back to me, which wasn&apos;t by design but turned out to be nice for this sort of workflow. The downside is that I&apos;ve now got a load of stuff to sift through - much of which is a &apos;good start&apos;, but none of it is likely to be usable without a good deal more effort. Here&apos;s a short-list of things I had it do:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Resurrect Fay Carson&apos;s work on the &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1295&quot;&gt;Menhir parser for odoc&lt;/a&gt;, pushed &lt;a href=&quot;https://github.com/jonludlam/odoc/tree/menhir-parser-rebased&quot;&gt;here&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Added some instrumentation to Odoc to do some performance experiments&lt;/li&gt;
&lt;li&gt;Ran some simple experiments to measure the impact of various pre-existing performance knobs/switches&lt;/li&gt;
&lt;li&gt;Resurrected an old patch of mine to &lt;a href=&quot;https://github.com/jonludlam/odoc/tree/parameterised-paths&quot;&gt;unify the two path representations&lt;/a&gt; in odoc to measure its effect on performance.&lt;/li&gt;
&lt;li&gt;Tested aggresively reuse of records if their fields don&apos;t change during compile/link&lt;/li&gt;
&lt;li&gt;Mixed up the &lt;a href=&quot;https://tangled.org/jon.recoil.org/odoc-scrollycode-extension&quot;&gt;scrollycode backend&lt;/a&gt; and the x-ocaml backend and stuck a playground on at each step&lt;/li&gt;
&lt;li&gt;Unified the oxcaml/ocaml branches of &lt;a href=&quot;https://tangled.org/jon.recoil.org/js_top_worker&quot;&gt;js_top_worker&lt;/a&gt; and x-ocaml via cppo&lt;/li&gt;
&lt;li&gt;Added oxcaml mode/layout annotations to odoc&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;OxCaml&lt;/h2&gt;
&lt;p&gt;I investigated the oxcaml docs build, which I had got working last week. Anil reported that it wasn&apos;t working for him, so I looked at the build I had and it definitely &lt;em&gt;was&lt;/em&gt; working. However, I was building on our machine Monteverde, which is a bit of a beast, so I checked the memory usage and it was enormous! I tried the build again on my 64 gig VM and it OOM&apos;d. I&apos;d noticed before that the &lt;code&gt;cmti&lt;/code&gt; files for base, in particular &lt;code&gt;base__Container.cmti&lt;/code&gt; were absolutely massive, and so had just assumed that the problem was that. Luke had also mentioned that some of the output from the template machinery was hidden. However, I had Claude look into this and it couldn&apos;t see any doc stop comments. So I asked it to look a little closer and figure out what was using all the memory. It took an unexpectedly large number of prods from me to finally figure out what was going on - it was to do with how odoc processes &lt;code&gt;includes&lt;/code&gt; - specifically an &lt;code&gt;include sig ... end&lt;/code&gt;. Essentially an include of that type ends up doubling the storage required of the signature. As the ppx_template extension does quite a lot of this, and in particular nests them, this ends up going exponential and this turned out to be the cause of most of the memory usage. With a fair bit more prodding by me, Claude and I eventually got to a solution, which I&apos;ll be upstreaming soon - the fix applies to OCaml as well as OxCaml, but it&apos;s this particularly pathalogical usage of includes that ppx_template uses where it&apos;ll make the most difference.&lt;/p&gt;
&lt;h2&gt;Odoc, plugins, JS and more&lt;/h2&gt;
&lt;p&gt;Teaser... I have a blog post coming soon with more on this. It&apos;s been a lot of fun, and should provide a decent inspiration for a roadmap for Odoc and online notebooks!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/02/weeknotes-2026-08.html" rel="alternate" title="Weeknotes weeks 7-8"/>
    <category term="weeknotes"/>
    <category term="ai"/>
    <category term="teaching"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/02/weeknotes-2026-06.html</id>
    <title type="text">Weeknotes for week 6</title>
    <updated>2026-02-09T00:00:00Z</updated>
    <published>2026-02-09T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Highlights:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://jon.ludl.am/experiments/day10-jtw/standalone/index.html&quot;&gt;day10 / javascript toplevels integration&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://jon.ludl.am/experiments/scrollycoder/&quot;&gt;Scrollycode experiments&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Oxmono&lt;/h2&gt;
&lt;p&gt;I spent some time on Anil&apos;s oxmono repo getting odoc to work correctly. It turned out that the bug I was working on last week was critically important for this - and that the bugfix was incomplete. One of the issues was to do with identifiers needing to be unique. For example, consider the following code:&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;module type S = sig
type t&lt;/p&gt;
&lt;p&gt;include sig
type t&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;val f : t -&amp;amp;gt; t
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;end with type t := t
end&lt;/x-ocaml&gt;
The problem here is that both definitions of `type t` have the same identifier, which causes problems when we move to and from the &apos;Component&apos; types. The solution was to introduce a &apos;dummy&apos; parent for the type defined within the include. This works because we never actually render the body of the include into HTML - we render the &lt;em&gt;expansion&lt;/em&gt;, which &lt;em&gt;doesn&apos;t&lt;/em&gt; have &lt;code&gt;type t&lt;/code&gt; in it, as it has been substituted out.&lt;/p&gt;
&lt;p&gt;The fix I made last week fixed the &lt;a href=&quot;/reference/odoc/odoc.loader/Odoc_loader/index.html&quot;&gt;loader&lt;/a&gt;, which reads in the &lt;code&gt;cmt&lt;/code&gt;/&lt;code&gt;cmti&lt;/code&gt; files produced by the compiler. There&apos;s one more place where we create these in the code - when we translate from the &lt;a href=&quot;/reference/odoc/odoc.xref2/Odoc_xref2/Component/index.html&quot;&gt;Component&lt;/a&gt; types back into &lt;a href=&quot;/reference/odoc/odoc.model/Odoc_model/Lang/index.html&quot;&gt;Lang&lt;/a&gt; types. I was a little curious about whether it was possible to make this happen, so I thought I&apos;d ask Claude to see if it could come up with a scenario where we&apos;d end up in this situation. This was a complete failure, which was a real disappointment to me, as doing this sort of thing is a quite tedious and annoying part of working on odoc.&lt;/p&gt;
&lt;p&gt;Meanwhile, I was running odoc on Anil&apos;s &lt;a href=&quot;https://github.com/avsm/oxmono&quot;&gt;oxmono&lt;/a&gt; repo, which was using &lt;a href=&quot;https://github.com/art-w&quot;&gt;art-w&lt;/a&gt;&apos;s &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1399&quot;&gt;PR to upstream oxcaml support&lt;/a&gt;. It was failing with an exception that was very familiar, so I pulled in the fix I&apos;d been working on, and that enabled it to get much further. However, it did subsequently fail with another slightly different exception. I had my suspicions at this point that it might be due to the other place, but I thought this again was a good opportunity to test Claude&apos;s debugging skills. However, this again was a complete failure. I spend quite a long time prodding it - at least 4 separate sessions - and it really didn&apos;t get anywhere close to a solution, despite knowing precisely that the commit we&apos;d made that had fixed the first problem. Two of the four times it ended up telling me that the oxcaml compiler was broken and suggesting that we create an issue!&lt;/p&gt;
&lt;p&gt;I&apos;m only very mildly disappointed in this - it&apos;s all quite subtle, and something I still end up scratching my head over sometimes, but it would have been wonderful to be able to offload this sort of work!&lt;/p&gt;
&lt;p&gt;In any case, the docs now all build on &lt;a href=&quot;https://github.com/jonludlam/oxmono/commit/2a53f6857d5b8849a73f5bb3e5244b9ac0f36708&quot;&gt;my fork of oxmono&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Docs CI&lt;/h2&gt;
&lt;p&gt;The fix I deployed last week for ocaml-docs-ci was taking forever to complete, so I ended up spending some time investigating this. The problem was happening during the &apos;prep&apos; phase, which is the first part of the pipeline where we simply build the package to be documented. This is supposed to work by building a graph of all inter-package dependencies across all of the solved packages, so we maximise sharing of built artefacts. Each &apos;prep&apos; job builds precisely one package by coping in the dependencies from previous prep jobs, then running &lt;a href=&quot;https://github.com/jonludlam/opamh&quot;&gt;opamh&lt;/a&gt; to fix up the metadata so that opam believes it has installed everything itself, then running opam to build the one package required. It was this last step that was going wrong, where it would decide that there had been upstream changes to the compiler itself, and rebuild &lt;em&gt;everything&lt;/em&gt;, so rather than a prep job taking a few seconds, it would take a few minutes.&lt;/p&gt;
&lt;p&gt;I was totally unable to repro this locally - everything build very quickly and just how it should have done. After much head-scratching I finally realised that the problem was somewhere in the caching. I think what&apos;s going on is that we dynamically build an opam repository to make the `opam install` command faster, and that repo contains only the packages that are required to build whatever it is we&apos;re building. Those opam files are cached by the docs CI server and passed to the build script as a base64-encoded gzipped tarball inline in the obuilder file (!). This should all be totally consistent as we&apos;re also caching all the builds - except for the compiler itself, which comes from the base docker image. This, of course, is the problem. The ocaml compiler opam files had been updated, and then when we reconstructed the opam repo with our cached opam files, opam noticed they had changed (gone &lt;em&gt;backwards&lt;/em&gt; in time!) and decided it needed to rebuild the compiler, and therefore &lt;em&gt;everything&lt;/em&gt; else. Clearing out the opam-files cache and restarting the builds fixed this entirely, and the full rebuild job completed after about 2 days. I flipped the switch on Saturday night and the docs are now fully up to date again. Phew!&lt;/p&gt;
&lt;h2&gt;day10 work&lt;/h2&gt;
&lt;p&gt;This was a fun week of large-scale building! I integrated day10 and odoc_driver and js_top_worker and x-ocaml and have now successfully got a docs-ci-like system that&apos;s able to build docs and toplevels that can coexist in the one HTML tree. I&apos;ve not got a full integrated demo yet, but you can see the test cases for this &lt;a href=&quot;https://jon.ludl.am/experiments/day10-jtw/standalone/index.html&quot;&gt;here&lt;/a&gt;. Be sure to take a look at the &apos;network&apos; tab in the browser dev tools to see what it&apos;s doing!&lt;/p&gt;
&lt;h2&gt;Scrollycode experiments&lt;/h2&gt;
&lt;p&gt;I&apos;ve long been a fan of &lt;a href=&quot;https://pomb.us/&quot;&gt;Rodrigo Pombo&apos;s&lt;/a&gt; work on &amp;quot;building tools for better code reading comprehension&amp;quot;, ever since first seeing his post &amp;quot;&lt;a href=&quot;https://pomb.us/build-your-own-react/&quot;&gt;Build your own React&lt;/a&gt;&amp;quot;. Claude is &lt;em&gt;fantastically good&lt;/em&gt; at doing this sort of thing, so I asked it to go and build me some simple OCaml-focused versions. We came up with 5 variations in the end - and they&apos;re all pretty neat! &lt;a href=&quot;https://jon.ludl.am/experiments/scrollycoder/&quot;&gt;take a look!&lt;/a&gt;. The best part of this was that it took me less than half-an-hour to get Claude to do all this.&lt;/p&gt;
&lt;h2&gt;Dune PR&lt;/h2&gt;
&lt;p&gt;I attended the bi-weekly dune dev meeting to talk about the first part of the dune PR - the bit that Paul Elliot did almost a year ago.&lt;/p&gt;
&lt;h2&gt;Coming week&lt;/h2&gt;
&lt;p&gt;So the clock is ticking on writing the exam questions for FoCS, so I&apos;ll need to be spending time this week on that.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/02/weeknotes-2026-06.html" rel="alternate" title="Weeknotes for week 6"/>
    <category term="weeknotes"/>
    <category term="notebooks"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/01/weeknotes-2026-04-05.html</id>
    <title type="text">Weeknotes for weeks 4-5</title>
    <updated>2026-01-30T00:00:00Z</updated>
    <published>2026-01-30T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;I&apos;ve been battling the seasonal illnesses this week, so I&apos;ve combined two weeknotes into one. Fortunately the &apos;flu doesn&apos;t hold Claude back!&lt;/p&gt;
&lt;p&gt;Probably the most interesting part of this is the &lt;a href=&quot;./#retrospective&quot;&gt;Retrospective&lt;/a&gt;, so make sure to read that bit.&lt;/p&gt;
&lt;h2&gt;The Last Two Weeks&lt;/h2&gt;
&lt;p&gt;As is becoming more and more apparent, the &lt;em&gt;breadth&lt;/em&gt; of what I&apos;m working on is ever expanding, powered by agentic AI. It&apos;s become so much more (cognitively) cheaper to have an idea and set an agent off investigating it that I&apos;ve been finding that I&apos;m working in parallel on far more things in a single week than I would have even six months ago. Here are some of the bigger headings though.&lt;/p&gt;
&lt;h3&gt;Monorepo excitement&lt;/h3&gt;
&lt;p&gt;We&apos;re currently experimenting with a new tool - &lt;a href=&quot;https://tangled.org/anil.recoil.org/monopam&quot;&gt;monopam&lt;/a&gt; to help develop across multiple OCaml libraries by using git subtrees to create a monorepo with all of the packages in. We then extract patches to the individual repos to push upstream. I&apos;ve been moving my development workflow from in-vscode-claude with careful permissions checking to running claude with `--dangerously-skip-permissions` in a container with the monorepo checked out. This has been a bit of a bumpy ride, with the tool evolving daily, but I&apos;m very much seeing the benefits of letting Claude just get on with things, given a strict enough early design and testing strategy, and using Anil&apos;s method of creating the interfaces first.&lt;/p&gt;
&lt;h4&gt;Odoc&lt;/h4&gt;
&lt;p&gt;I also did quite a bit related to odoc these 2 weeks, split over improving functionality and bugfixing.&lt;/p&gt;
&lt;h5&gt;Plugins&lt;/h5&gt;
&lt;p&gt;Getting Claude to run with all of the monorepo libraries implicitly requires that they&apos;re well documented, as looking at the source to figure out how to use them exhausts the context window pretty rapidly. Odoc&apos;s main focus has been on getting the expansions and referencing correct, and while we&apos;ve made progress on the actual content markup, introducing &lt;a href=&quot;https://ocaml.github.io/odoc/odoc/odoc_for_authors.html#media&quot;&gt;media tags&lt;/a&gt; for example, there&apos;s still a good distance to go.&lt;/p&gt;
&lt;p&gt;Using the plugins mechanism I &lt;a href=&quot;/blog/2026/01/weeknotes-2026-03.html&quot;&gt;wrote about last week&lt;/a&gt;, I&apos;ve made a plugin interface for odoc and implemented a few plugins. Initially I was just going to support &apos;custom tags&apos; but it occurred to me that rendering code blocks could also be done in this way. So I&apos;ve made a few. Two custom tag plugins:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-admonition-extension/&quot;&gt;odoc-admonition-extension&lt;/a&gt; - styled callout blocks for notes, warnings, tips. Note that we are intending to make this more first-class - there&apos;s a &lt;a href=&quot;https://hackmd.io/ETSOAmetTI-E3vrDk3Bfrw&quot;&gt;design out there&lt;/a&gt;. This was just a convenient way to test the feature!&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-rfc-extension/&quot;&gt;odoc-rfc-extension&lt;/a&gt; - links to IETF RFC documents
and 3 code block plugins:&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-msc-extension/&quot;&gt;odoc-msc-extension&lt;/a&gt; - Message Sequence Charts&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-mermaid-extension/&quot;&gt;odoc-mermaid-extension&lt;/a&gt; - Mermaid diagrams (flowcharts, sequence diagrams, etc.)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;/reference/odoc-dot-extension/&quot;&gt;odoc-dot-extension&lt;/a&gt; - Graphviz/DOT diagrams
The module signatures relevant to the plugins are documented in &lt;a href=&quot;/reference/nox-odoc/nox-odoc.extension_api/Odoc_extension_api/index.html&quot;&gt;&lt;code&gt;Odoc_extension_api&lt;/code&gt;&lt;/a&gt; and the plugins each have to implement an interface described in &lt;a href=&quot;/reference/nox-odoc/nox-odoc.extension_api/Odoc_extension_api/module/type/Code_Block_Extension/index.html&quot;&gt;&lt;code&gt;Odoc_extension_api.Code_Block_Extension&lt;/code&gt;&lt;/a&gt; or &lt;a href=&quot;/reference/nox-odoc/nox-odoc.extension_api/Odoc_extension_api/module/type/Extension/index.html&quot;&gt;&lt;code&gt;Odoc_extension_api.Extension&lt;/code&gt;&lt;/a&gt; for custom tags.&lt;/li&gt;
&lt;/ul&gt;
&lt;h5&gt;Bugfixing&lt;/h5&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/lukemaurer&quot;&gt;Luke Maurer&lt;/a&gt; at Jane Street pointed out that they&apos;re still suffering from yet another repro of &lt;a href=&quot;https://github.com/ocaml/odoc/issues/930&quot;&gt;issue 930&lt;/a&gt; at Jane Street. I&apos;d worked on this &lt;a href=&quot;/blog/2025/09/odoc-bugs.html&quot;&gt;back in September&lt;/a&gt; but turns out I hadn&apos;t actually made a PR, so I tidied up the branch and &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1400&quot;&gt;made a PR&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;Docs CI&lt;/h3&gt;
&lt;p&gt;Docs CI has been fixed and is even now rebuilding all of the docs for ocaml.org. I&apos;ve added in the &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci/commit/c6231fa383820b4c700aaa1e72107536b1872112&quot;&gt;handling of `post &amp;amp; with-doc`&lt;/a&gt; in place of x-extra-doc-deps, so we should be able to use either mechanism now. The idea is to deprecate x-extra-doc-deps soon though. Somehow despite an explicit button to press to update the epoch symlinks, it got updated anyway and broke most of the docs on ocaml.org. Fortunately &lt;a href=&quot;https://discuss.ocaml.org/t/is-caqti-doc-missing/17741/5&quot;&gt;someone noticed&lt;/a&gt; and posted on discuss and so I switched it back.&lt;/p&gt;
&lt;p&gt;Unfortunately, it seemed to be taking a long time to build the docs - at time of writing it&apos;s now Friday, and the CI jobs have been running since Tuesday. In that time, it&apos;s only managed to build about 6500 packages, a long way short of the 16,000 or so that I expect a full build will produce. Looking through the logs, it seems that some change to opam is causing it to sometime rebuild the entire opam universe when it should only be building 1 package. For example, in a job that should be building just `tezos-protocol-004-Pt24m4xi`, it installs all of the prebuilt dependencies, then runs `opamh` to try to convince opam that everything is all set up to just run the build step for the package we want. Unfortunately the logs show the following:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;The following actions will be performed:
=== recompile 178 packages
  - recompile aches                            1.1.0              [uses ocaml]
  - recompile aches-lwt                        1.1.0              [uses ocaml]
...
  - recompile mtime                            2.1.0              [uses ocaml]
  - recompile ocaml                            4.14.2             [upstream or system changes]
  - recompile ocaml-compiler-libs              v0.12.4            [uses ocaml]
...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where it seems opam has decided that something has changed enough for it to want to recompile the `ocaml` package, and therefore &lt;em&gt;everything&lt;/em&gt; in the entire opam switch! So this job took 12 minutes instead of 21 seconds, which was the time required to finally build the `tezos-protocol` package.&lt;/p&gt;
&lt;h3&gt;Day10 and docs&lt;/h3&gt;
&lt;p&gt;In closely related news, &lt;a href=&quot;https://tunbury.org/&quot;&gt;mtelver&apos;s&lt;/a&gt; day10 project looked precisely the right shape for building docs - in fact it shares its architecture and some components with the docs CI. So I asked Claude to take a look and see what it would take, and discovered that it doesn&apos;t take very much! We have a Really Big Machine here at the CL that was temporarily underused; and by Really Big I mean 768 cores and 3TB of RAM. So, how long could building all of the docs for all of the packages possibly take? Well, it takes 5 hours 40 mins. And I was only using roughly a third of the machine. Nice!&lt;/p&gt;
&lt;p&gt;So should I push on with fixing ocaml-docs-ci and figure out why it&apos;s rebuilding everything all the time? Or should I forge ahead with day10 and turn it into a proper CI system as opposed to a slightly flakey bespoke thing I have to handhold through a build? This is next week&apos;s problem.&lt;/p&gt;
&lt;h3&gt;JS toplevels&lt;/h3&gt;
&lt;p&gt;Something I keep coming back to is javascript toplevels. I&apos;d really like to be able to be able to host JS toplevels on ocaml.org for each different version of each different package. This is something I&apos;ve worked on on-and-off for a long time now, and several fixes to help have been merged to various projects along the way. The tricky thing is to not put a massive load onto ocaml.org with this, so we need to be efficient. That means firstly having a single toplevel js file with all of the logic in but none of the libraries, and then dynamically loading libraries as we need them. Also we can save some bandwidth by not immediately sending all of the cmi files, as these can be faulted in as necessary too. So once again I&apos;ve got Claude on the task, and things are honestly looking pretty hopeful now. I&apos;ve got 2 demos:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://jon.ludl.am/experiments/findlibish/&quot;&gt;Dynamic library loading&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://jon.ludl.am/experiments/multi-universe-demo/&quot;&gt;Multi-version support&lt;/a&gt;
In both cases, make sure you take a look at the network tab to see it dynamically loading only what it needs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Retrospective&lt;/h2&gt;
&lt;h3&gt;Autonomous Claude&lt;/h3&gt;
&lt;p&gt;The power of sending Claude off to do some work can be immense. However, it does mean investing time up front telling it precisely what problem you&apos;re trying to solve, what approach to take, finer details on how you want it done, and how you can tell if it&apos;s working when it finishes. A &apos;failure mode&apos; I&apos;ve been experiencing is when I end up in a long, drawn out real time interaction, especially if that&apos;s happening with 2 projects simultaneously - and by &apos;failure&apos; I really mean just &apos;slow&apos;. Ideally what would be going on is for all of my agents to be getting on with whatever task they&apos;ve been allocated without bothering me for more details. For Claude to have to ask me a question has much more latency involved than it just getting on with things, especially if I don&apos;t notice it immediately.&lt;/p&gt;
&lt;h3&gt;When to Stop&lt;/h3&gt;
&lt;p&gt;The &apos;finishing criteria&apos; are important - many times this week I&apos;ve had Claude tell me it&apos;s finished something, having verified that it&apos;s passing all the tests, only for me to take a look to find that it&apos;s very obviously broken. As quite a few things recently have involved the web, I&apos;ve put Playwright into all of my devcontainers, and told Claude to use it to verify things are working. This has been working pretty well, so I&apos;ll be adding it to my prompts. It&apos;s not too dissimilar to what we used to call &apos;pre-flight checks&apos; back in the Citrix days.&lt;/p&gt;
&lt;h3&gt;Containers vs accounts&lt;/h3&gt;
&lt;p&gt;I&apos;ve been running everything with `--dangerously-ignore-permissions` in containers, and while the outcome is amazing, the containers bit has been a bit of a headache. Next week I&apos;ll be trialling the idea of just giving the agents their own account (non-admin!) on my servers, their own github account, tangled account and so on, and just treating them more like I would if I had a real colleague. It&apos;s always slightly alarming to see my own name on the output of the bots, assigning me (or sometimes someone else (!!)) copyright over code I&apos;ve never seen. This is, of course, a whole other pandora&apos;s box that I really don&apos;t want to open right now - but I think the point is that I&apos;ll feel a lot more comfortable if the commits are all by `Jon&apos;s Agent &amp;lt;jon+claude@recoil.org&amp;gt;` rather than by me!&lt;/p&gt;
&lt;h3&gt;Deciding next steps&lt;/h3&gt;
&lt;p&gt;The question of whether I should fix up ocaml-docs-ci or improve the day10 solution requires a bit of thought. In fact, it requires a bit of a gap analysis between the two. This isn&apos;t something I&apos;ve asked Claude to do before, so I&apos;ll try that and see how it turns out. I&apos;ll be asking it to be &amp;quot;scientific&amp;quot; in its approach, coming up with hypotheses and verifying them - for which I think I&apos;ll need to give it a platform on which it can perform experiments. This is a bit trickier with ocaml-docs-ci than day10 as day10 runs entirely on any given linux computer, whereas ocaml-docs-ci needs ocurrent workers and a routable ssh server. I&apos;ll report on the outcome of this next week!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/01/weeknotes-2026-04-05.html" rel="alternate" title="Weeknotes for weeks 4-5"/>
    <category term="weeknotes"/>
    <category term="odoc"/>
    <category term="ai"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2026/01/weeknotes-2026-03.html</id>
    <title type="text">Weeknotes for week 3</title>
    <updated>2026-01-19T00:00:00Z</updated>
    <published>2026-01-19T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;First week back of 2026! Let&apos;s write some terse weeknotes.&lt;/p&gt;
&lt;h2&gt;Projects&lt;/h2&gt;
&lt;h3&gt;Dune odoc rules&lt;/h3&gt;
&lt;p&gt;Last thing I did last year was to push the new rules for odoc 3. This week, Anil handed me an excellent opportunity to test the rules on the monorepo containing his &lt;a href=&quot;https://anil.recoil.org/notes/aoah-2025&quot;&gt;AOAH&lt;/a&gt; projects. Claude tends to actually write ocamldoc-formatted comments, so this is really useful to test the rules. I&apos;ve &lt;a href=&quot;https://github.com/jonludlam/dune/tree/odoc-v3-rules-3.21&quot;&gt;rebased the commits&lt;/a&gt; on the just-released Dune 3.21 and we&apos;ve been trying them out. There were a few things to fix:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;More careful &lt;a href=&quot;https://github.com/jonludlam/dune/commit/25158eabf0c3cac2826e16ce590b4bd4d7c09818&quot;&gt;dependency tracking&lt;/a&gt; during the compile phase - this particularly affected the &lt;code&gt;@doc&lt;/code&gt; target, which was pulling in unnecessary dependencies. Most of these dependencies were compiling just fine, but one - Anstrom - is slightly odd in that the opam install of Angstrom installs a META file that references libraries that aren&apos;t in the dependencies of its opam package. This is a backward-compatibility hack that was implemented when the Anstrom package was split into several in order to manage the dependencies better.&lt;/li&gt;
&lt;li&gt;A similar issue happens with eio, where the documentation of the package depends upon &lt;code&gt;bigstring&lt;/code&gt;, which isn&apos;t in eio&apos;s dependencies. This is entirely intentional - the extra doc dependencies is stated in the opam file with a &lt;code&gt;x-extra-doc-deps&lt;/code&gt; field. However, &lt;code&gt;opam install&lt;/code&gt; totally ignores this field (quite reasonably), and so a simple install gives you an opam repo whose docs can&apos;t be built. Once again, this broke &lt;code&gt;dune build @doc&lt;/code&gt; unnecessarily, but the fix was &lt;a href=&quot;https://github.com/jonludlam/dune/commit/2afe046cf4290d7a83b5f2c5646e3391ca94b630&quot;&gt;relatively simple&lt;/a&gt;. The &lt;em&gt;real&lt;/em&gt; fix here is to not use &lt;code&gt;x-extra-doc-deps&lt;/code&gt;, but switch to using a &lt;em&gt;real&lt;/em&gt; dependency, but marked with &lt;code&gt;with-doc&lt;/code&gt; and &lt;code&gt;post&lt;/code&gt; if it would otherwise introduce a circular dependency. That way, an &lt;code&gt;opam install --with-doc&lt;/code&gt; &lt;em&gt;would&lt;/em&gt; install the extra dependency.&lt;/li&gt;
&lt;li&gt;Over the Christmas break, &lt;a href=&quot;https://discuss.ocaml.org/u/tbrk&quot;&gt;tbrk&lt;/a&gt; posted &lt;a href=&quot;https://discuss.ocaml.org/t/odoc-index-for-multiple-packages-inter-package-links-and-local-global-sidebar/17652&quot;&gt;on discuss&lt;/a&gt; a question about building docs, for which my dune branch was a partial answer. One feature he was requesting though was the ability to use a custom top-level index. It&apos;s a useful feature that&apos;s implemented in &lt;code&gt;odoc_driver&lt;/code&gt; so I&apos;ve &lt;a href=&quot;https://github.com/jonludlam/dune/commit/efecdee1b36b7e47906e7c64b7496a1fc7954a2d&quot;&gt;added it&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;More sensible &lt;a href=&quot;https://github.com/jonludlam/dune/commit/039eb3d2a3e9d28f8b195905f43839daf5ce8c21&quot;&gt;default link scope&lt;/a&gt;. By default, documentation references in the &lt;code&gt;mli&lt;/code&gt; files of a library can link to any other library in the package. However, by default it wasn&apos;t possible to link to the dependencies of another library, unless it happened to be a dependency of your own library. Similarly, the package-wide mld files could only reference the modules in the package&apos;s libraries, not to the dependencies. This seems overly cautious, as we can be sure that if we&apos;ve managed to build the libraries then their dependencies are installed, and if there are any module name conflicts, we can resolve them via the &lt;code&gt;/&amp;lt;lib&amp;gt;/Module&lt;/code&gt; syntax.&lt;/li&gt;
&lt;li&gt;Lastly, implementations of virtual libraries &lt;a href=&quot;https://github.com/jonludlam/dune/commit/12f9ecbd4888444c2d359049a914ffb4827912f9&quot;&gt;need to be skipped&lt;/a&gt; as they&apos;ve all got the same docs (as they share mli files), and the rules as they were causing Dune to crash with a &amp;quot;Conflicting implementations&amp;quot; error.
I&apos;ve also rebased the PR onto latest &lt;code&gt;main&lt;/code&gt;, but I&apos;ve not yet put these patches there, which I&apos;ll need to do for the PR to be mergable. For now, the 3.21 branch is successfully building the docs for the monorepo.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;OCaml Docs CI&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/jmid&quot;&gt;Jan Midtgaard&lt;/a&gt; noticed over xmas that the Docs CI &lt;a href=&quot;https://github.com/ocaml/ocaml.org/issues/3437&quot;&gt;was broken&lt;/a&gt; and submitted &lt;a href=&quot;https://github.com/jonludlam/opamh/pull/1&quot;&gt;a fix&lt;/a&gt;. I&apos;ve therefore been poking &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci&quot;&gt;ocaml-docs-ci&lt;/a&gt; to get the fix incorporated and into production. I almost immediately hit the issue that &lt;code&gt;odoc_driver&lt;/code&gt; now breaks for the exact same reason. I couldn&apos;t quite understand how &lt;code&gt;opam-format&lt;/code&gt; &lt;a href=&quot;https://github.com/ocaml/opam-repository/pull/28978&quot;&gt;had been merged&lt;/a&gt; to &lt;code&gt;opam-repository&lt;/code&gt; without someone noticing that it had broken &lt;code&gt;odoc_driver&lt;/code&gt;, but it turned out that it &lt;em&gt;had&lt;/em&gt; been noticed, but on a &lt;a href=&quot;https://github.com/ocaml/opam-repository/pull/28877&quot;&gt;beta release&lt;/a&gt;. The fix to docs ci was to install &lt;code&gt;odoc_driver&lt;/code&gt; from opam rather than &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci/blob/81ca17c7b7a2f47ca571b1d6bc866720cebef136/src/lib/config.ml#L226&quot;&gt;pinning directly&lt;/a&gt; to a github hash, especially if that hash happens to be the hash of the released version!&lt;/p&gt;
&lt;p&gt;While I&apos;m working on docs CI, I thought it&apos;s probably also a good idea to move over to the &lt;code&gt;with-doc &amp;amp; post&lt;/code&gt; suggestion from above, so we&apos;re ready for when packages start to use that. This is now being tested, and hopefully we&apos;ll have the CI back up and running early next week.&lt;/p&gt;
&lt;h3&gt;Better styling for odoc&lt;/h3&gt;
&lt;p&gt;I&apos;ve done very little to the styling of odoc since I took maintainership way back in 2019 or so. It&apos;s a bit dated, and there are some annoying usability issues, so I thought it&apos;s a good opportunity to vibe-code a nice new frontend for it. Rather than hack directly on the HTML generator of odoc, this seemed to be a good opportunity to test the JSON output from the new Dune rules, so I asked Claude to make me a static site generator that read in the JSON files and spat out some nicely styled HTML. This worked like a charm, and the results are &lt;a href=&quot;https://jon.ludl.am/experiments/vibe-coded-odoc-frontend/&quot;&gt;here&lt;/a&gt;. Next steps are to see what it would take to get the native odoc output looking more like that.&lt;/p&gt;
&lt;h3&gt;Custom tags in odoc&lt;/h3&gt;
&lt;p&gt;One of the themes of Anil&apos;s &lt;a href=&quot;https://anil.recoil.org/notes/aoah-2025&quot;&gt;AOAH&lt;/a&gt; coding spree was that many libraries were implementations of RFCs. In many places in the docs, there are links to relevant sections of the RFCs. It&apos;d be nice in future to be able to validate that we&apos;ve covered all of the parts of the RFCs, so making the links a little more parsable seemed like a good idea. In fact, it seemed that this might be a perfect use for custom tags - a feature that was present in ocamldoc that odoc has yet to implement.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/art-w&quot;&gt;Arthur Wendling&lt;/a&gt; then pointed me at dune&apos;s &lt;a href=&quot;https://dune.readthedocs.io/en/stable/reference/dune/plugin.html&quot;&gt;plugin system&lt;/a&gt;, which seemed just the ticket as a way to implement this. It&apos;s really nice, taking all of the hard work out of creating OCaml plugins, so I&apos;ve now got &lt;a href=&quot;https://github.com/jonludlam/odoc/tree/extension-plugins&quot;&gt;an extension-plugins branch&lt;/a&gt; that implements this. It allows you to add support to odoc for tags like &lt;code&gt;@rfc&lt;/code&gt; which generate custom HTML, markdown or any other backend, can include links in their bodies, and can add custom headers to the web page, and custom files to be output by &lt;code&gt;odoc support-files&lt;/code&gt;. It looks like this should &amp;quot;just work&amp;quot; and no further changes to the dune rules are needed - though I need to actually test this out.&lt;/p&gt;
&lt;h3&gt;Day10 and docs&lt;/h3&gt;
&lt;p&gt;I&apos;ve &lt;a href=&quot;/blog/2025/09/build-ids-for-day10.html&quot;&gt;written about&lt;/a&gt; &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark&apos;s&lt;/a&gt; day10 project before. It&apos;s a tool to very rapidly build odoc packages mainly in order to test that they build correctly. An obvious extension would be to use this to then build the docs for those packages, as the way we do this requires the packages to be built first. This would be a replacement for the Docs CI that I talked about above, though there&apos;s considerable work to do before it&apos;s fully-featured enough to be a viable alternative. It seemed like a good time to experiment with this though, so I set up one of Anil&apos;s &lt;a href=&quot;https://anil.recoil.org/notes/ocaml-claude-dev&quot;&gt;devcontainers&lt;/a&gt;, gave Claude some instructions on what to do, took the safety belt off, and let him hack away! Previously most of my interactions with Claude had been via the vscode plugin, so using the terminal interface was a bit of a different experience. I&apos;m fairly certain though that I&apos;m going to switch everything over to working this way, as letting Claude just get on with things without having to OK every step is a far more efficient way to work - especially when you&apos;re not that concerned with the actual code being produced. This has been mostly a good experience, though Claude does sometimes go off in rather odd directions. At one point there was a network error with a dependency while trying to build odoc_driver, so it decided that it should have a fallback mechanism that executed odoc directly. I told it &lt;em&gt;NEVER&lt;/em&gt; to replace functionality in odoc_driver, so it rolled this back, but a few hours later in then did exactly the same thing again.&lt;/p&gt;
&lt;h3&gt;Misc other stuff&lt;/h3&gt;
&lt;p&gt;A few other things too - &lt;a href=&quot;https://github.com/jonludlam/odoc/commit/59037341cd53d8734a5874f7af2b728b5be70035&quot;&gt;improving the &lt;code&gt;--warn-error&lt;/code&gt; logic in odoc&lt;/a&gt;, and one of its &lt;a href=&quot;https://github.com/jonludlam/odoc/commit/9d18feff5eda543652c6749062750de6e5bb4d6e&quot;&gt;error messages&lt;/a&gt;, improving the build of this website so I can iterate on it more quickly, fixing up some of my self-hosted services like my tangled knot, and other bits and bobs.&lt;/p&gt;
&lt;h2&gt;Reflections&lt;/h2&gt;
&lt;p&gt;I think the most important thing this week has been the slightly eye-opening benefits of using Claude outside of the context of VSCode. I suspect I&apos;ll be doing much more of my work this way in future. There&apos;s also a good chance I&apos;ll have to upgrade my subscription from the $100-per-month to the $200 one...&lt;/p&gt;
&lt;h2&gt;Next week&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Start of term tutorial meetings&lt;/li&gt;
&lt;li&gt;Sherldoc in monopam-myspace&lt;/li&gt;
&lt;li&gt;Get ocaml-docs-ci deployed and working&lt;/li&gt;
&lt;li&gt;Update the Dune PR&lt;/li&gt;
&lt;li&gt;Integrate the custom-tags and website generator into monopam-myspace&lt;/li&gt;
&lt;li&gt;Unleash Claude on my js-top-worker repo&lt;/li&gt;
&lt;/ul&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2026/01/weeknotes-2026-03.html" rel="alternate" title="Weeknotes for week 3"/>
    <category term="weeknotes"/>
    <category term="odoc"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/12/claude-and-dune.html</id>
    <title type="text">Claude and Dune</title>
    <updated>2025-12-18T00:00:00Z</updated>
    <published>2025-12-18T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Back in March of this year we released &lt;a href=&quot;https://ocaml.github.io/odoc/odoc/index.html&quot;&gt;odoc 3.0.0&lt;/a&gt;, a major new version of the OCaml documentation generator. It had a whole load of &lt;a href=&quot;https://discuss.ocaml.org/t/ann-odoc-3-beta-release/16043&quot;&gt;new features&lt;/a&gt;, many of which came with new demands on the build system driving it. We decided when working on it to build a new driver for odoc so that we could adjust it as we were building the new features, and this driver is now used to &lt;a href=&quot;/blog/2025/07/odoc-3-live-on-ocaml-org.html&quot;&gt;build the documentation&lt;/a&gt; that appears on &lt;a href=&quot;https://ocaml.org/p/base/latest/doc/index.html&quot;&gt;ocaml.org&lt;/a&gt;. However, it was always the plan to integrate the new features into &lt;a href=&quot;https://dune.build&quot;&gt;Dune&lt;/a&gt; so that everyone could just run &lt;code&gt;dune build @doc&lt;/code&gt; and be able to use all of the new odoc 3 features.&lt;/p&gt;
&lt;p&gt;So over the last few weeks I have been wrestling with getting Claude to update the odoc rules in Dune to support &lt;em&gt;some&lt;/em&gt; of the new features of odoc v3. What began as a background experiment during a lecture series has turned into a multi-week effort to turn mostly-working code into a clean, reviewable patch. AI-developed software is clearly going to be a big part of our future, and Anil is showing us all the way with his &lt;a href=&quot;https://anil.recoil.org/notes/aoah-2025-1&quot;&gt;Advent of Agentic Humps&lt;/a&gt; by building &lt;em&gt;new&lt;/em&gt; software, but upstreaming AI-generated changes to an existing, well established code base &lt;a href=&quot;https://github.com/ocaml/ocaml/pull/14369&quot;&gt;hasn&apos;t got off to a good start&lt;/a&gt; in the OCaml community, so I wanted to be extra careful to get this right.&lt;/p&gt;
&lt;h3&gt;Claude as a protyping tool&lt;/h3&gt;
&lt;p&gt;The initial progress was pretty amazing, despite my initial worries that the dune code-base would be &lt;a href=&quot;https://github.com/ocaml/dune/pull/12529&quot;&gt;too large and subtle&lt;/a&gt; for an LLM to be able to make workable changes. In order to get going, first I had it look at several bits of example code:&lt;/p&gt;
&lt;p&gt;1. &lt;a href=&quot;https://github.com/ocaml/dune/blob/3.20.2/src/dune_rules/odoc.ml&quot;&gt;dune_rules/odoc.ml&lt;/a&gt; - this is the current home of the odoc rules in dune. It&apos;s local-only, meaning it only builds the docs for the current package in isolation, so no resolution of links to stdlib, other packages or libraries.&lt;/p&gt;
&lt;p&gt;2. &lt;a href=&quot;https://github.com/ocaml/dune/blob/3.20.2/src/dune_rules/odoc_new.ml&quot;&gt;dune_rules/odoc_new.ml&lt;/a&gt; - these are the rules for odoc v2, which allow you to build the docs for your package plus all of the dependencies. I wrote this mostly myself some time ago. It does a pretty poor job of caching, error reporting, and has none of the odoc v3 features like assets, source rendering, hierarchical docs, better errors and so on.&lt;/p&gt;
&lt;p&gt;3. &lt;a href=&quot;https://github.com/ocaml/odoc/tree/d8460cdaa2b91a03434a9a045d673703b7fabfb2/src/driver&quot;&gt;odoc_driver&lt;/a&gt; - this is the driver we wrote when building odoc v3. It&apos;s fully featured, but not at all incremental, and actually external to the dune codebase. It&apos;s the reference implementation that&apos;s used to build the docs that appear on &lt;a href=&quot;https://ocaml.org/p/base/latest/doc/index.html&quot;&gt;ocaml.org&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Armed with these three code-bases, I asked Claude to synthesise a new incremental version of the odoc rules for dune that has some of the features of &lt;code&gt;odoc_driver&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;The working prototype&lt;/h3&gt;
&lt;p&gt;Claude quickly produced a prototype that actually compiled and generated documentation. At that stage I was not interested in the quality of the generated source; I only needed to know whether Claude could navigate Dune&apos;s codebase and produce something that &lt;strong&gt;works&lt;/strong&gt;. I let the prototype evolve incrementally, adding in new features one at a time, for example, fixing the error reporting so that you only get warned about documentation errors that you can actually fix.&lt;/p&gt;
&lt;p&gt;When the lectures finished, it turned out I had something that was pretty useful to me, and had a good chance to be useful to others too. So I opened up my editor and had a look through what had been produced, at this point hoping that a little bit of polishing should be enough - after all, it &lt;em&gt;was&lt;/em&gt; working!&lt;/p&gt;
&lt;p&gt;It was dreadful.&lt;/p&gt;
&lt;p&gt;There were long, rambling functions, code duplication, bad comments, it was unstructured, with repeated-but-slightly-different chunks all over the place. It wasn&apos;t just bad on one length scale - it was bad from the large-scale organisation of the code down to small scale baffling weirdnesses on one line. The more I looked, the more bonkers it appeared. But it did &lt;em&gt;work&lt;/em&gt;! So I thought I&apos;d get Claude to clean up its own messes.&lt;/p&gt;
&lt;h3&gt;The clean-up&lt;/h3&gt;
&lt;p&gt;I resolved that I would continue to let Claude do &lt;em&gt;all&lt;/em&gt; of the editing, and not do &lt;em&gt;any&lt;/em&gt; myself, and so thus began the more frustrating part of this adventure! I ended up giving a mix of very specific instructions: &amp;quot;move this code here&amp;quot;, &amp;quot;factorize out this functionality&amp;quot;, &amp;quot;rename this function&amp;quot;, and sometimes more general ones: &amp;quot;Remove any comments that don&apos;t add anything of value&amp;quot;, or &amp;quot;Think of a better way to do this&amp;quot;. The constant was that I needed to be looking over each change that it did, because while most of them were pretty good, there were still a few, even with the very explicit instructions, where it messed up. From the very broad, where at one point it told me &amp;quot;I&apos;ll remove this code to create odoc files for external dependencies, as they&apos;re installed by opam&amp;quot;, which isn&apos;t true, down to the very small - for example, it produced the following:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;let lib_names = deps.Odoc_config.libraries in
if List.is_empty lib_names
then Memo.return []
else Memo.List.filter_map lib_names ~f:(fun lib_name -&amp;gt; Lib.DB.find lib_db lib_name)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where it has come up with a totally redundant check for the empty list.&lt;/p&gt;
&lt;p&gt;It was at this point where it became frustrating, because although it&apos;s almost magical that Claude can do what it does in the time it does, this fact of having to keep a constant eye on it meant the the tens-of-seconds to minutes delay in between it doing something meant I ended up either twiddling my thumbs for long periods of time, or getting started on some other task and forgetting to come back to Claude, sometimes for hours!&lt;/p&gt;
&lt;h3&gt;OCaml is &lt;strong&gt;not&lt;/strong&gt; the problem&lt;/h3&gt;
&lt;p&gt;One part that particularly impressed, and also quite surprised me, was with its knowledge of OCaml. In particular, I had at one point two different types representing the &apos;target&apos; - either a library or a package - and a &apos;kind&apos; - either a module or a page. Now pages can only be associated with package targets, and modules can only be associated with libraries, but these two values were distinct, so there was a fair bit of code pattern matching invalid combinations and either throwing exceptions or picking some random value, depending on the whims of Claude&apos;s context. I bravely suggested it think of a better way to represent this, maybe using GADTs, and it did indeed come up with a pretty nice refactoring of the types:&lt;/p&gt;
&lt;p&gt;Before:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;type target =
  | Lib of Package.Name.t * Lib.t
  | Pkg of Package.Name.t

type artifact_kind =
  | Module of
      { visible : bool
      ; module_name : Module_name.t
      ; archive : string (* Which archive the module belongs to *)
      }
  | Page of
      { name : string
      ; pkg_libs : Lib.t list
      }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;After:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;(* Artifact data types *)
type page = { name : string; pkg_libs : Lib.t list }

type mod_ =
  { visible : bool
  ; module_name : Module_name.t
  ; archive : string (* Which archive the module belongs to *)
  }

type _ target =
  | Lib : Package.Name.t * Lib.t -&amp;gt; mod_ target
  | Pkg : Package.Name.t -&amp;gt; page target

type artifact_kind =
  | Module : mod_ * mod_ target -&amp;gt; artifact_kind
  | Page : page * page target -&amp;gt; artifact_kind
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This refactoring immediately removed a whole swathe of invalid combinations, making the code both safer and clearer. It&apos;s quite clear that Claude had no trouble understanding how GADTs work in OCaml, quite happily also using some existentials to pack them into lists and so on.&lt;/p&gt;
&lt;h3&gt;Odd behaviours&lt;/h3&gt;
&lt;p&gt;Sometimes Claude just went a little bit bananas. One annoyance that &lt;em&gt;repeatedly&lt;/em&gt; occurred was that it would forget how to build and test the dune executable, despite clear instructions in &lt;code&gt;Claude.md&lt;/code&gt;. Most of the time when it went wrong it would build dune, execute &lt;code&gt;dune clean&lt;/code&gt;, then try to run the dune binary that it had just removed with the &lt;code&gt;clean&lt;/code&gt;. Sometimes it would decide to use the bootstrap binary instead, which isn&apos;t rebuilt on every change, sometimes it would run the switch-installed dune binary, and on one occasion it tried to run &lt;code&gt;./configure &amp;amp;&amp;amp; make&lt;/code&gt;!&lt;/p&gt;
&lt;p&gt;It would usually figure out eventually what the right thing to do was, but when you&apos;re waiting for it to complete so you can check what it&apos;s done these sorts of delays got a bit frustrating.&lt;/p&gt;
&lt;h3&gt;Reflections&lt;/h3&gt;
&lt;p&gt;At one point, I ran out of Claude credits (despite paying $100 a month or so), at about 6:20pm one evening, and it told me that I needed to wait until 7pm to carry on. I&apos;d just got to the point when I needed to write a short bit of code rather than refactoring what was already there, and I realised that while it would take me maybe 10 mins, it would take Claude maybe 10 seconds. Now, it could just be that it was the end of a long day and I was running out of steam, but I was content to switch focus elsewhere for a bit to wait for my credits to reset before carrying on! The point being that for the small implementation that I was after, it would be possible for me to get Claude to do it, and to eyeball the result to make sure it was OK in less time than I would have been able to do it myself. But I absolutely wouldn&apos;t have trusted Claude to do it in an upstreamable way &lt;strong&gt;without&lt;/strong&gt; looking at the result.&lt;/p&gt;
&lt;p&gt;Overall, It&apos;s clear that Claude will be an incredibly useful tool for working with software. It&apos;s unbelievably good at jumping into a new code-base and figuring things out quickly, but less good at producing high-quality code that can be directly submitted upstream (yet?) - at least, not that &lt;strong&gt;I&lt;/strong&gt; would be comfortable submitting anyway. However, I think it&apos;s still a bit of an open question as to what the quality bar &lt;em&gt;should&lt;/em&gt; be. If it builds correctly, passes the tests, looks &lt;em&gt;broadly&lt;/em&gt; sensible and isn&apos;t on the critical path for performance, how much should we care about the line-to-line quality? &lt;strong&gt;I&lt;/strong&gt; certainly care, but am I being old fashioned?&lt;/p&gt;
&lt;p&gt;I&apos;ve submitted a &lt;a href=&quot;https://github.com/ocaml/dune/pull/12995&quot;&gt;PR with these changes&lt;/a&gt; for review, and we&apos;ll see what happens there. I ended up squashing all of the commits into one, as the intermediate steps are very likely not useful. However, for historical interest, the branch on which I did most of the work is &lt;a href=&quot;https://github.com/ocaml/dune/compare/main...jonludlam:dune:odoc3-global-sidebar&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/12/claude-and-dune.html" rel="alternate" title="Claude and Dune"/>
    <category term="odoc"/>
    <category term="ai"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/12/an-svg-is-all-you-need.html</id>
    <title type="text">An SVG is all you need</title>
    <updated>2025-12-09T00:00:00Z</updated>
    <published>2025-12-09T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;SVGs are pretty cool - vector graphics in a simple XML format. They are supported on just about every device and platform, are crisp on every display, and can have embedded scripts in to make them interactive. They&apos;re &lt;a href=&quot;https://www.youtube.com/watch?v=4laPOtTRteI&quot;&gt;way more capable&lt;/a&gt; than many people realise, and I think we can capitalise on some of that unrealised potential.&lt;/p&gt;
&lt;p&gt;Anil&apos;s recent post &lt;a href=&quot;https://anil.recoil.org/notes/principles-for-collective-knowledge&quot;&gt;Four Ps for Building Massive Collective Knowledge Systems&lt;/a&gt; got me thinking about the permanence of the experimentation that underlies our scientific papers. In my idealistic vision of how scientific publishing should work, each paper would be accompanied by a fully interactive environment where the reader could explore the data, rerun the experiments, tweak the parameters, and see how the results changed. Obviously we can&apos;t do this in the general case - some experiments are just too expensive or time-consuming to rerun on demand. But for many papers, especially in computer science, this is entirely feasible.&lt;/p&gt;
&lt;p&gt;That line of thought reminded me of a project I tackled about 20 years ago as a post-doc in the Department of Plant Sciences here in Cambridge. I was writing a paper on &lt;a href=&quot;https://royalsocietypublishing.org/rsif/article/9/70/949/173/Applications-of-percolation-theory-to-fungal&quot;&gt;synergy in fungal networks&lt;/a&gt; and built a tiny SVG visualisation tool that let readers wander through the raw data captured from a real fungal network growing in a petri dish. I dug it up recently and was surprised (and delighted) to see that it still works perfectly in modern browsers - even though the original “cover page” suggested Firefox 1.5 or the Adobe SVG plug-in (!). Give it a spin; click the &apos;forward&apos;, &apos;back&apos; and other buttons below the petri dish!&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./fungus.svg&quot; alt=&quot;fungus.svg&quot; &gt;
And that, dear reader, is literally all you need. A completely self-contained SVG file can either fetch data from a versioned repository or embed the data directly, as the example does. It can process that data, generate visualisations, and render knobs and sliders for interactive exploration. No server-side magic required - everything runs client-side in the browser, served by a plain static web server, and very easily to share.&lt;/p&gt;
&lt;p&gt;How does it fit in with Anil&apos;s four Ps?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Permanence: SVGs can be assigned DOIs just like papers, blog posts, or datasets. The fact that the above SVG still works after two decades is a testament to the durability of the format.&lt;/li&gt;
&lt;li&gt;Provenance: Because SVG is plain text, it plays nicely with version control systems such as Git. When an SVG pulls in external data, the same provenance-tracking strategies Anil describes for datasets apply here as well.&lt;/li&gt;
&lt;li&gt;Permission: Once again, with the separation between the processing in the SVG and that data that it works on, the same permissioning models apply as for data in general.&lt;/li&gt;
&lt;li&gt;Placement: SVGs are &lt;em&gt;inherently&lt;/em&gt; spatial; it&apos;s very easy, for example, to make beautiful &lt;a href=&quot;https://stephanwagner.me/coding/blog/create-world-map-charts-with-svgmap#svgMapDemoGDP&quot;&gt;world maps&lt;/a&gt; with SVG.
The SVG above is only a visualisation tool for data; it doesn&apos;t really do any processing, but it certainly &lt;em&gt;could&lt;/em&gt;. The biggest change that&apos;s happened over the 20 years since I wrote this is the &lt;em&gt;massive&lt;/em&gt; increase in the computation power available in the browser. If would be entirely feasible to implement the entire data analysis pipeline for that paper in an SVG today, probably without even spinning up the fans on my laptop!&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So this is yet another tool in our ongoing effort to be able to effortlessly share and remix our work - added to the pile of Jupyter notebooks, &lt;a href=&quot;https://digitalflapjack.com/blog/marimo/&quot;&gt;Marimo botebooks&lt;/a&gt;, the &lt;a href=&quot;https://slipshow.readthedocs.io/en/stable/&quot;&gt;slipshow&lt;/a&gt;/&lt;a href=&quot;https://github.com/art-w/x-ocaml/&quot;&gt;x-ocaml&lt;/a&gt; &lt;a href=&quot;/blog/2025/11/foundations-of-computer-science.html&quot;&gt;combination&lt;/a&gt;, &lt;a href=&quot;https://patrick.sirref.org/weekly-2025-w45/index.xml&quot;&gt;Patrick&apos;s take&lt;/a&gt; on Jon Sterling&apos;s &lt;a href=&quot;https://sr.ht/~jonsterling/forester/&quot;&gt;Forester&lt;/a&gt;, my own &lt;a href=&quot;/notebooks/&quot;&gt;notebooks&lt;/a&gt;, and many others - and this is a subset of what we&apos;re using just in our own group!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/12/an-svg-is-all-you-need.html" rel="alternate" title="An SVG is all you need"/>
    <category term="notebooks"/>
    <category term="plugins"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/11/foundations-of-computer-science.html</id>
    <title type="text">Foundations of Computer Science</title>
    <updated>2025-11-14T00:00:00Z</updated>
    <published>2025-11-14T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;I recently completed lecturing the course &lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/&quot;&gt;&amp;quot;Foundations of Computer Science&amp;quot;&lt;/a&gt; to our newly arrived first-year computer scientists here at &lt;a href=&quot;https://www.cam.ac.uk&quot;&gt;Cambridge&lt;/a&gt;. This is the first time I&apos;ve lectured this course, taking over from &lt;a href=&quot;https://anil.recoil.org/&quot;&gt;Anil&lt;/a&gt; while he&apos;s on sabbatical. Although I was very nervous indeed about it, I ended up really enjoying the experience - and I hope the students did too! This post is a little brain dump of my thoughts on how it went and how we might improve it for next year.&lt;/p&gt;
&lt;h2&gt;Course Overview&lt;/h2&gt;
&lt;p&gt;The course is 12 lectures long and has been lectured in a similar way since I myself was an undergraduate here, way back in 1996. There have been a few changes, not least of which is that back then it was in Standard ML rather than OCaml, but the core material has remained largely the same: lists, recursive functions, trees, higher-order functions, search and finally mutability. There are no prerequisites for the course, although all students have at least a maths A-level (or equivalent), and almost all of them have done some programming before, though the experience varies widely. Very few have done any functional programming, and even fewer have written any OCaml before.&lt;/p&gt;
&lt;p&gt;The notes for the course are distributed both in hard copy and also as an &lt;a href=&quot;https://github.com/ocamllabs/focs-notebooks/blob/main/1A%20Foundations%20of%20Computer%20Science.ipynb&quot;&gt;interactive Jupyter Notebook&lt;/a&gt;, which we host on our &lt;a href=&quot;https://hub.cl.cam.ac.uk&quot;&gt;JupyterHub server&lt;/a&gt; that I maintain. The idea is that the students can read through the notes and then play around with the code examples directly in the notebook. I don&apos;t encourage them or give them time to do much &lt;em&gt;during&lt;/em&gt; the lectures - not that I think this is a terrible idea, but it&apos;s a struggle to fit all the material in otherwise! The notes are pretty closely coupled to the lectures, organised into 11 chapters that correspond to the first 11 lectures, with exercises at the end of each chapter that are intended to be covered in the supervisions. We also have some assessed exercises - &amp;quot;Ticks&amp;quot; - that the students complete in their own time using the JupyterHub server using &lt;a href=&quot;https://github.com/jupyter/nbgrader&quot;&gt;nbgrader&lt;/a&gt;. They are automatically assessed in a very transparent way; each &amp;quot;tick&amp;quot; is a Jupyter notebook with editable answer cells and read-only test cells. Overall we&apos;re aiming for the students not to &lt;em&gt;have&lt;/em&gt; to install OCaml locally at all, though I hope many of them will choose to do so anyway.&lt;/p&gt;
&lt;p&gt;While I didn&apos;t want them playing around with the notebook during the lectures, I do, however, try to get them to interact by getting them to answer questions. It&apos;s pretty intimidating to stick your head above the parapet like this, so as an incentive I rewarded those that answered (rightly or wrongly) with some of the excellent stickers that Tarides has printed over the years. Everybody loves stickers!&lt;/p&gt;
&lt;p&gt;The questions I asked varied quite a lot in their difficulty, and many were in the first few minutes of each lecture, where I had a short &apos;warm-up&apos; where we recapped the contents of the previous lecture. These warm-ups were strongly suggested by Anil, and as well as reminding everyone of where we left off, they also gave me a bit of feedback on the things that the students found challenging.&lt;/p&gt;
&lt;p&gt;One entertaining aspect is that during the first lecture I do actually encourage them to at least log on to the JupyterHub server, mostly to get them used to the idea of trying it. The entertaining part is that our server isn&apos;t particularly big and beefy, and so with 130 students all trying to log on at once, it invariably caves in under the load. At this point in the lecture I ssh to the server and run btop/htop and we watch it die in real time!&lt;/p&gt;
&lt;h2&gt;What changed this year&lt;/h2&gt;
&lt;p&gt;During the lectures themselves, rather than use Keynote or PowerPoint for the slides, I decided to try using &lt;a href=&quot;https://slipshow.readthedocs.io/en/stable/&quot;&gt;Slipshow&lt;/a&gt;, augmented with &lt;a href=&quot;https://github.com/art-w/x-ocaml&quot;&gt;x-ocaml&lt;/a&gt; to embed executable OCaml code snippets. I&apos;m very happy with how this worked out. I was able to prepare both working and broken snippets, modify them live during the lecture, and things like type-on-hover was very useful. In a few lectures where we were discussing big-O notation, I was able to run code on different input sizes and really demonstrate the big difference in run-time of certain algorithms. After the lectures, I posted the slides onto the course website so that students can refer back to them, and they can also try out the live code snippets directly in the slides.&lt;/p&gt;
&lt;p&gt;Both Slipshow and x-ocaml are still quite young projects, so it was inevitable that there were a few rough edges, and in fact the interaction of the two revealed the biggest problem: that when you use the &apos;speaker-view&apos; mode of Slipshow, where you have a separate window with notes and the current slide, the x-ocaml widgets are effectively independent in the two windows, so updating in one doesn&apos;t update in the other. &lt;a href=&quot;https://choum.net/panglesd/&quot;&gt;Paul-Elliot&lt;/a&gt;, the author of Slipshow, had already got a potential fix for this in the works when I spoke to him about it, so hopefully next time I use this I&apos;ll be able to have speaker notes on screen, instead of hand-written index cards! The x-ocaml project is a lot smaller than Slipshow, so I was able to use Claude to help me add functionality I needed, such as being able to programmatically highlight sections of the code.&lt;/p&gt;
&lt;p&gt;Another new thing I tried this year was to go over &apos;tracing&apos; of execution to help the students understand how programs run. We&apos;ve always taught reduction steps in the course, which works well as it&apos;s only the last lecture where we introduce mutability, but it can quickly become unwieldy, and it can be challenging to do this all by hand. Tracing a function tells the runtime to log when function calls and returns happen, so you just need to call the function on your desired input, and you get a fully automatic trace of the execution. As it&apos;s only function calls and returns, it doesn&apos;t tell the full story, but alongside the handwritten reduction, it can help reassure students that they&apos;re on the right track. I ended up writing up a trace of a particularly complicated lazy-list evaluation using Slipshow and x-ocaml, which I posted &lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/interleave_explanation.html&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Thoughts for next year&lt;/h2&gt;
&lt;p&gt;Overall I&apos;m very happy with how the course went this year, though in some ways it did feel a little bit like the course finished just when it had started to get to the good stuff! There&apos;s a Tripos review process going on at the moment, so maybe we&apos;ll get to expand this course a bit in future years.&lt;/p&gt;
&lt;p&gt;While the Slipshow+x-ocaml combination worked well, the fact that we ended up with two separate systems for executing OCaml wasn&apos;t ideal. I think it&apos;d be a really nice project to investigate just how far we can push x-ocaml / Slipshow / some other web technology to have a true &amp;quot;serverless&amp;quot; experience so we can ditch the JupyterHub server entirely. By caching the x-ocaml &apos;execution&apos; web worker in the browser, we could have a system that works fully offline, removing an annoyingly failure-prone single point of failure. Of course, we&apos;d still need some way to do the assessed exercises, but that&apos;s a small point in a much larger problem: we really can&apos;t continue to ignore how LLMs are impacting the way that students are approaching these exercises in both positive and negative ways. To answer this properly, we need to think hard about what the purpose of these exercises is and look around to see what our &lt;a href=&quot;https://eecs.iisc.ac.in/people/prof-viraj-kumar/&quot;&gt;colleagues&lt;/a&gt; are doing &lt;a href=&quot;https://dl.acm.org/doi/10.1145/3724363.3729100&quot;&gt;in this space&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The slide decks themselves are fully open and available on the &lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/&quot;&gt;course website&lt;/a&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture1/lecture1.html&quot;&gt;Introduction&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture2/lecture2.html&quot;&gt;Recursion and Complexity&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture3/lecture3.html&quot;&gt;Lists and Polymorphism&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture4/lecture4.html&quot;&gt;More Lists and Making Change&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture5/lecture5.html&quot;&gt;Sorting&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture6/lecture6.html&quot;&gt;Datatypes and Trees&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture7/lecture7.html&quot;&gt;Dictionaries and Functional Arrays&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture8/lecture8.html&quot;&gt;Currying&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture9/lecture9.html&quot;&gt;Sequences, or Lazy Lists&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture10/lecture10.html&quot;&gt;Search&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture11/lecture11.html&quot;&gt;Procedural Programming&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2526/FoundsCS/lecture12/lecture12.html&quot;&gt;Recap and Real World Use!&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/11/foundations-of-computer-science.html" rel="alternate" title="Foundations of Computer Science"/>
    <category term="teaching"/>
    <category term="ocaml"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/09/caching-opam-solutions2.html</id>
    <title type="text">Caching opam solutions - part 2</title>
    <updated>2025-09-23T00:00:00Z</updated>
    <published>2025-09-23T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Some results from the &lt;a href=&quot;/blog/2025/09/caching-opam-solutions.html&quot;&gt;previous post&lt;/a&gt;. This time I&apos;ve run day10 on 144 or so commits from opam-repository to see how well the cache performs. The results are quite interesting.&lt;/p&gt;
&lt;p&gt;First let&apos;s talk about the &amp;quot;examination map&amp;quot;. This is a map from package name to a list of other packages whose solutions should be recalculated if the package in question is altered. It&apos;s built by first looking at the packages that the solver asks about during the solution for a package, and then taking &lt;em&gt;all&lt;/em&gt; of the solutions, and &apos;inverting&apos; the map, so for example, if both packages &apos;a&apos; and &apos;b&apos; ask about package &apos;c&apos; during their solutions, then altering &apos;c&apos; means that the solutions for both &apos;a&apos; and &apos;b&apos; need to be recalculated. The examination map entry for &apos;c&apos; would then be &lt;code&gt;&apos;a&apos;; &apos;b&apos;&lt;/code&gt;. We can plot the histogram of the sizes of each entry in the examination map:&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./examination_map_histogram.svg&quot; alt=&quot;Package Examiner Distribution Histogram&quot; &gt;
Some interesting features from these data:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The most common number of observers is 1, meaning that the package is not involved in the solution of any other package. There are approximately 2000 such packages.&lt;/li&gt;
&lt;li&gt;Most (~80%) of packages have fewer than 100 observers. This means that if we alter one of these packages, we only need to recalculate the solutions for fewer than 100 other packages.&lt;/li&gt;
&lt;li&gt;A &lt;em&gt;very&lt;/em&gt; small number of packages are observed in all 4,400 solutions. This is actually a bit artificial, as the solver adds the ocaml-compiler package as an input to all solves to ensure we get the correct compiler version. There&apos;s another way to do this which would avoid this particular problem.&lt;/li&gt;
&lt;li&gt;A small number of packages have a very large number of observers, around 3800. This mostly corresponds with &lt;code&gt;dune&lt;/code&gt; and its dependencies and associated packages. There are around 350 such packages, and any change to these means we need to recalcuate most of the solutions.
This last point doesn&apos;t mean that we actually &lt;em&gt;recompile&lt;/em&gt; 3,800 packages, just that we need to recalcualte the solution, which might then lead to a cache hit of the layer and no actual compilation. However, recalculating the solutions of all of the packages takes (on my computer) around 10,000 seconds, or roughly 5 minutes of wall-clock time as I&apos;ve got 32 threads.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;However, if the package that&apos;s changes &lt;em&gt;isn&apos;t&lt;/em&gt; one of those 350 packages, then the number of solutions that need to be recalculated is dramatically reduced. I ran the logic over the last few weeks of commits to opam-repository, from commit &lt;code&gt;109398e2fd61803126becd398df0f1eabc9f3ca2&lt;/code&gt; of the 10th September up until commit &lt;code&gt;3f21ebe342ce440d9c9142ffe1185d8e5a326085&lt;/code&gt; from the 22nd. In this time there were 144 commits (counting only those from &lt;code&gt;git log --first-parent&lt;/code&gt;). Of these, only 4 resulted in a full resolve - the first commit, since obviously we have no cache at that point, the &lt;a href=&quot;https://github.com/ocaml/opam-repository/commit/40283204789e7116e1c99466de902cd565d121cf&quot;&gt;release of OCaml 5.4.0 beta2&lt;/a&gt; by &lt;a href=&quot;https://perso.quaesituri.org/florian.angeletti/&quot;&gt;Florian Angeletti&lt;/a&gt;, a fix of &lt;a href=&quot;https://github.com/ocaml/opam-repository/commit/6ef6813522b6ea29933f6451236a1639bdbaec61&quot;&gt;ocaml-base-compiler for MSVC&lt;/a&gt; by &lt;a href=&quot;https://www.dra27.uk/blog/&quot;&gt;David&lt;/a&gt; and a fix for &lt;a href=&quot;https://github.com/ocaml/opam-repository/commit/d141887ab0b4fc0836ad0787f1f806585a260bc8&quot;&gt;BER-OCaml&lt;/a&gt; by &lt;a href=&quot;https://www.cl.cam.ac.uk/~jdy22/&quot;&gt;Jeremy Yallop&lt;/a&gt;. Then 25 commits resulted in recalculating solutions for 3800 packages as they hit dune-adjacent packages, 5 commits resulted in recalculating between 100 and 300 packages and the remaining 110 commits resulted in recalculating fewer than 100 packages, the majority of which resulted in recalculating fewer than 5 packages.&lt;/p&gt;
&lt;p&gt;Overall, at a rough estimate, this means that over this period, using this caching strategy gave us a 5x speedup in the solver!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/09/caching-opam-solutions2.html" rel="alternate" title="Caching opam solutions - part 2"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/09/odoc-bugs.html</id>
    <title type="text">Odoc bugs</title>
    <updated>2025-09-22T00:00:00Z</updated>
    <published>2025-09-22T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;This post is a brief write-up of a couple of bugs in odoc that I&apos;ve been working on over the past 2 weeks. I was convinced at the start of this that I was actually fixing one bug, but although they both had the same backtrace and similar immediate causes, they&apos;re actually quite different. They both involve &lt;em&gt;expansion&lt;/em&gt;, which is the process that odoc uses to work out the contents of a module from its expression - what allows you to see the contents of a module such as &lt;code&gt;module M = Map.Make(String)&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;Bug 930: inline destructive substitutions&lt;/h3&gt;
&lt;p&gt;Bug #930 in odoc is about a substitution problem:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type S1 = sig
  type t0
  type &apos;a t := unit

  val x : t0 t
end

module type S2 = sig
  type t (* must be the same name as [S1.t] *)

  include S1 with type t0 := t
end

module type S3 = sig
  type t1

  include S2 with type t := t1
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;which when processed by odoc 2.4 throws an exception:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;odoc: internal error, uncaught exception:
      Invalid_argument(&amp;quot;List.fold_left2&amp;quot;)
      Raised at Stdlib.invalid_arg in file &amp;quot;stdlib.ml&amp;quot;, line 33, characters 20-45
      Called from Odoc_xref2__Subst.type_expr in file &amp;quot;subst.ml&amp;quot;, line 598, characters 21-59
      Called from Odoc_xref2__Subst.value in file &amp;quot;subst.ml&amp;quot; (inlined), line 842, characters 19-38
      Called from Odoc_xref2__Subst.apply_sig_map.inner.(fun) in file &amp;quot;subst.ml&amp;quot;, line 1089, characters 19-52
      Called from Odoc_xref2__Component.Delayed.get in file &amp;quot;component.ml&amp;quot; (inlined), line 55, characters 16-22
      Called from Odoc_xref2__Lang_of.signature_items.inner in file &amp;quot;lang_of.ml&amp;quot;, line 438, characters 16-39
      Called from Odoc_xref2__Lang_of.signature in file &amp;quot;lang_of.ml&amp;quot; (inlined), line 466, characters 12-43
      Called from Odoc_xref2__Lang_of.include_ in file &amp;quot;lang_of.ml&amp;quot;, line 641, characters 18-69
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The key thing here is that definition of &lt;code&gt;&apos;a t&lt;/code&gt; in &lt;code&gt;S1&lt;/code&gt; - a destructive substituion. If you type this code into an OCaml toplevel, you will see that the signature of &lt;code&gt;S1&lt;/code&gt; is:&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot; run-on=&quot;load&quot;&gt;module type S1 = sig
type t0
type &apos;a t := unit&lt;/p&gt;
&lt;p&gt;val x : t0 t
end&lt;/x-ocaml&gt;
where the substitution has clearly taken place. In contrast, odoc takes the position that the use of these inline destructive substitutions is to make the code easier to understand, and so it tries to keep them in the signature rather than simply apply them and present the resulting signature. So when rendering &lt;code&gt;S1&lt;/code&gt; we end up with:&lt;/p&gt;
&lt;div class=&quot;inset&quot; style=&quot;border: 1px solid var(--pre-border-color); padding: 10px; border-radius: 5px&quot;&gt;
&lt;a id=&quot;module-type-S1&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;h2&gt;Module type &lt;code&gt;&lt;span&gt;S1&lt;/span&gt;&lt;/code&gt;&lt;/h2&gt;
&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec type anchored&quot; id=&quot;type-t0&quot;&gt;&lt;a href=&quot;#type-t0&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; t0&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec type subst anchored&quot; id=&quot;type-t&quot;&gt;&lt;a href=&quot;#type-t&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; &lt;span&gt;&apos;a t&lt;/span&gt;&lt;/span&gt;&lt;span&gt; := unit&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec value anchored&quot; id=&quot;val-x&quot;&gt;&lt;a href=&quot;#val-x&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;val&lt;/span&gt; x : &lt;span&gt;&lt;a href=&quot;#type-t0&quot;&gt;t0&lt;/a&gt; &lt;a href=&quot;#type-t&quot;&gt;t&lt;/a&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;The reported problem is a failure with a stack trace while processing &lt;code&gt;S3&lt;/code&gt;, but upon looking closely the real problem has happened when expanding &lt;code&gt;S2&lt;/code&gt;. What happens is that we have a type &lt;code&gt;t&lt;/code&gt; defined in &lt;code&gt;S2&lt;/code&gt; and a type &lt;code&gt;t&lt;/code&gt; that will later be substituted away that comes from the inclusion of &lt;code&gt;S1&lt;/code&gt;. The rendered signature of &lt;code&gt;S2&lt;/code&gt; is:&lt;/p&gt;
&lt;div class=&quot;inset&quot; style=&quot;border: 1px solid var(--pre-border-color); padding: 10px; padding-right:30px; border-radius: 5px&quot;&gt;
&lt;a id=&quot;module-type-S2&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;h2&gt;Module type &lt;code&gt;&lt;span&gt;S2&lt;/span&gt;&lt;/code&gt;&lt;/h2&gt;
&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec type anchored&quot; id=&quot;type-s2-t&quot;&gt;&lt;a href=&quot;#type-s2-t&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; t&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;odoc-include&quot;&gt;&lt;details open=&quot;open&quot;&gt;&lt;summary class=&quot;spec include&quot;&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;include&lt;/span&gt; &lt;a href=&quot;#module-type-S1&quot;&gt;S1&lt;/a&gt; &lt;span class=&quot;keyword&quot;&gt;with&lt;/span&gt; &lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; &lt;a href=&quot;#type-t0&quot;&gt;t0&lt;/a&gt; := &lt;a href=&quot;#type-s2-t&quot;&gt;t&lt;/a&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/summary&gt;&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec type subst anchored&quot; id=&quot;type-s2-t&quot;&gt;&lt;a href=&quot;#type-s2-t&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; &lt;span&gt;&apos;a t&lt;/span&gt;&lt;/span&gt;&lt;span&gt; := unit&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec value anchored&quot; id=&quot;val-x&quot;&gt;&lt;a href=&quot;#val-x&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;val&lt;/span&gt; x : &lt;span&gt;&lt;a href=&quot;#type-s2-t&quot;&gt;t&lt;/a&gt; &lt;a href=&quot;#type-s2-t&quot;&gt;t&lt;/a&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;/details&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;where the type of &lt;code&gt;x&lt;/code&gt; is now &lt;code&gt;t t&lt;/code&gt;, which is clearly incorrect. The problem is that odoc assumes that type names are unique within a signature (modulo shadowing, which isn&apos;t quite what&apos;s going on here), but in this signature there are two definitions of &lt;code&gt;type t&lt;/code&gt;, one of which is parameterised and one is not. At this point nothing fatal has happened, but when we try to process &lt;code&gt;S3&lt;/code&gt; the substitution code gets very confused by these different arities and &lt;code&gt;List.fold_left2&lt;/code&gt; throws the above exception.&lt;/p&gt;
&lt;p&gt;The fix I&apos;m trialling for this is that when we&apos;re including a signature that contains an inline destructive substitution, we will perform that substitution when the expansion of the include is done. This means that the rendered signature of &lt;code&gt;S1&lt;/code&gt; will be just the same as before, but the rendered signature of &lt;code&gt;S2&lt;/code&gt; will now be:&lt;/p&gt;
&lt;div class=&quot;inset&quot; style=&quot;border: 1px solid var(--pre-border-color); padding: 10px; padding-right:30px; border-radius: 5px&quot;&gt;
&lt;a id=&quot;module-type-newS2&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;h2&gt;Module type &lt;code&gt;&lt;span&gt;S2&lt;/span&gt;&lt;/code&gt;&lt;/h2&gt;
&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec type anchored&quot; id=&quot;type-s2new-t&quot;&gt;&lt;a href=&quot;#type-s2new-t&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; t&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;odoc-include&quot;&gt;&lt;details open=&quot;open&quot;&gt;&lt;summary class=&quot;spec include&quot;&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;include&lt;/span&gt; &lt;a href=&quot;#module-type-S1&quot;&gt;S1&lt;/a&gt; &lt;span class=&quot;keyword&quot;&gt;with&lt;/span&gt; &lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;type&lt;/span&gt; &lt;a href=&quot;#type-t0&quot;&gt;t0&lt;/a&gt; := &lt;a href=&quot;#type-s2new-t&quot;&gt;t&lt;/a&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/summary&gt;&lt;div class=&quot;odoc-spec&quot;&gt;&lt;div class=&quot;spec value anchored&quot; id=&quot;val-x&quot;&gt;&lt;a href=&quot;#val-x&quot; class=&quot;anchor&quot;&gt;&lt;/a&gt;&lt;code&gt;&lt;span&gt;&lt;span class=&quot;keyword&quot;&gt;val&lt;/span&gt; x : unit&lt;/span&gt;&lt;/code&gt;&lt;/div&gt;&lt;/div&gt;&lt;/details&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;where the type of &lt;code&gt;x&lt;/code&gt; is now simply &lt;code&gt;unit&lt;/code&gt;, which is what OCaml itself thinks, happily! I think this strikes the balance between keeping the substitutions visible for clarity where they are originally defined, but when including them elsewhere we simply see the resulting signature.&lt;/p&gt;
&lt;h3&gt;Bug #1385: Exception raised during compilation&lt;/h3&gt;
&lt;p&gt;The second bug has the identical backtrace, indicating a problem with arities. However, the repro case for this one does not involve any inline destructive substitution, though it does involve destructive substitution at the module expression level:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type Creators_base = sig
  type (&apos;a, _, _) t
  type (_, _, _) concat

  include sig
      type (&apos;a, &apos;b, &apos;c) t

      val concat : ((&apos;a, &apos;p1, &apos;p2) t, &apos;p1, &apos;p2) concat -&amp;gt; (&apos;a, &apos;p1, &apos;p2) t
    end
    with type (&apos;a, &apos;b, &apos;c) t := (&apos;a, &apos;b, &apos;c) t
end

module type S0_with_creators_base = sig
  type t

  include Creators_base with type (&apos;a, _, _) t := t and type (&apos;a, _, _) concat := t
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;There&apos;s quite a lot of type parameters flying around here, so the first step was to try to simplify this as much as possible while still getting the exception. I got it down to:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type Creators_base = sig
  type &apos;a t
  type _ concat

  include sig
      type &apos;a t

      val concat : &apos;a concat -&amp;gt; &apos;a t
    end
    with type &apos;a t := &apos;a t
end

module type S0_with_creators_base = sig
  type t

  include Creators_base with type _ t := t with type _ concat := t
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;which still throws the same exception. So, what&apos;s going on here? Fundamentally, it&apos;s a similar issue to the first bug, just caused in a different way, in that once again we&apos;ll end up with a signature that has two definitions of &lt;code&gt;type t&lt;/code&gt; with different arities. In this case, the problem occurs during the expansion of &lt;code&gt;S0_with_creators_base&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This is the intermediate expansion of &lt;code&gt;Creators_base&lt;/code&gt; that odoc calculates:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type S0_with_creators_base = sig
  type t

  include Creators_base with type _ t := t with type _ concat := t (*
    
    The expansion as calculated by odoc is:

      include sig
        type &apos;a t
        val concat : t -&amp;gt; &apos;a t
      end
  *)
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;What&apos;s happened here is during the calculation of the body of the include, odoc has taken the signature of &lt;code&gt;Creators_base&lt;/code&gt; and has its two type definitions both replaced with &lt;code&gt;type t&lt;/code&gt; (with no parameters). However, since the &lt;code&gt;type t&lt;/code&gt; in the body of the include is defined in that signature, that one wasn&apos;t replaced. So we end up with the type of &lt;code&gt;concat&lt;/code&gt; being &lt;code&gt;t -&amp;gt; &apos;a t&lt;/code&gt;, which looks very odd! At this point though, odoc knows very well that they&apos;re different types. However, when odoc converts this signature back into the datatype that represents the expansions, it loses that information and we end up with the two types mixed up. We then go on to process this signature, the mixup of the arities causes the failure.&lt;/p&gt;
&lt;p&gt;There are several independent fixes that we can make here. Firstly we can make sure that we don&apos;t mix up the types. This we can do because we can distinguish between items that are declared within the signature of the include&apos;s declaration and those that come from the outer context. We don&apos;t have to do this for the expansion of the include as OCaml&apos;s type system means that there can&apos;t be two types of the same name in the resulting signature. We never actually render any signature that occurs within the body of an include, so this doesn&apos;t actually make any difference to the output.&lt;/p&gt;
&lt;p&gt;The second fix is to make sure that we only calculate the expansion of the include once. Currently the bug happens because we try to re-calculate the expansion of the &lt;code&gt;include sig ... end&lt;/code&gt; expression, even though we calculated it during the processing of &lt;code&gt;S0_with_creators_base&lt;/code&gt;. What we should do instead is apply the substitutions to the expansion of that calculated include, which would end up with the same result. This isn&apos;t a perfect solution though, as there are occasions when we have to recalculate the signature anyway.&lt;/p&gt;
&lt;p&gt;The third fix is - and this takes a little care to parse - to ensure that we never actually try to process the items within a signature within a &amp;quot;with&amp;quot; expression within a module-type expression. Before diving into the &apos;why&apos; of this, let&apos;s first explain how Odoc represents module-type expressions.&lt;/p&gt;
&lt;p&gt;Internally, we have &lt;a href=&quot;/reference/odoc/odoc.model/Odoc_model/Lang/ModuleType/index.html#type-expr&quot;&gt;a datatype&lt;/a&gt; that represents module expressions, which looks like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;type expr =
    | Path of path_t
    | Signature of Signature.t
    | Functor of FunctorParameter.t * expr
    | With of with_t
    | TypeOf of typeof_t
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now, each of the arguments to these constructors might contain an expansion of the expression that Odoc will calculate. For example, the definition of &lt;a href=&quot;/reference/odoc/odoc.model/Odoc_model/Lang/ModuleType/index.html#type-path_t&quot;&gt;path_t&lt;/a&gt; is:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;type path_t = {
    p_expansion : simple_expansion option;
    p_path : Paths.Path.ModuleType.t;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and this expansion is initially &lt;code&gt;None&lt;/code&gt; and then filled in by Odoc in order to render the expansion in the HTML. In the case of a &lt;code&gt;With&lt;/code&gt; expression, the &lt;a href=&quot;/reference/odoc/odoc.model/Odoc_model/Lang/ModuleType/index.html#type-with_t&quot;&gt;with_t&lt;/a&gt; type is:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;type with_t = {
    w_substitutions : substitution list;
    w_expansion : simple_expansion option;
    w_expr : U.expr;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;here you can see that the &lt;code&gt;With&lt;/code&gt; expression contains another module expression, as a &lt;code&gt;with&lt;/code&gt; expression operates on another module type. Early during Odoc&apos;s development, this simply was another `ModuleType.expr`, but we had a couple of bugs where we ended up calculating expansions for these inner expressions, which was all very wasteful as we only ever rendered the &amp;quot;outer&amp;quot; expansion. So we changed this to be a &lt;a href=&quot;/reference/odoc/odoc.model/Odoc_model/Lang/ModuleType/U/index.html#type-expr&quot;&gt;U.expr&lt;/a&gt;, which is an &amp;quot;unexpanded&amp;quot; module type expression, and is very similar to the main expression above, but without the expansions and also with the functor case, as we can&apos;t have functors inside a &amp;quot;with&amp;quot; expression.&lt;/p&gt;
&lt;p&gt;These &amp;quot;unexpanded&amp;quot; expressions still contain signatures though, so aren&apos;t &lt;em&gt;completely&lt;/em&gt; unexpanded, and it&apos;s &lt;em&gt;these&lt;/em&gt; signatures that we should avoid processing.&lt;/p&gt;
&lt;p&gt;So, what I expected to be just one bug when I started looking at this turned out to be two related issues, and a total of four different fixes!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/09/odoc-bugs.html" rel="alternate" title="Odoc bugs"/>
    <category term="odoc"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/09/caching-opam-solutions.html</id>
    <title type="text">Caching opam solutions</title>
    <updated>2025-09-09T00:00:00Z</updated>
    <published>2025-09-09T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;The &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci&quot;&gt;ocaml-docs-ci&lt;/a&gt; system works by watching opam-repository for changes, and then when it notices a new package it performs an opam solve and builds the package, a prerequisite for building the documentation. In order to give the docs some stability, as the docs may well &lt;a href=&quot;/blog/2025/04/semantic-versioning-is-hard.html&quot;&gt;depend upon your dependencies&lt;/a&gt;, we currently cache the solve results so that a package will always be built with the same set of dependencies, even if a new version of one of those dependencies has been released.&lt;/p&gt;
&lt;p&gt;The downside to this is that as time goes on, the number of distinct universes that we build increases, and docs get more and more out of date. So it&apos;s not necessarily the best thing to do, though it does mean we minimise the amount of time spent solving.&lt;/p&gt;
&lt;p&gt;The alternative approach is that on every commit to opam-repository we could resolve for all packages and use the latest, greatest solution to build the docs. Using this approach we would maximise the sharing of builds and keep the total amount of required storage steadier. Of course, this would mean solving for every package on every commit to opam-repository, even if we didn&apos;t end up rebuilding all of them due to the way that the cache works.&lt;/p&gt;
&lt;p&gt;One possibility that might be worth investigating is to cache the solutions - but then Leon Bambrick &lt;a href=&quot;https://twitter.com/secretGeek/status/7269997868&quot;&gt;advises us&lt;/a&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-quote&quot;&gt;There are 2 hard problems in computer science: cache invalidation,
naming things, and off-by-1 errors.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and indeed it&apos;s not obvious what the best approach to cache invalidation is here. A sledgehammer approach would be to hook into the solver and note what questions it asks of opam-repository and record the responses. If any of these change, then it&apos;s safe to say that we need to recalculate. I had a quick look at this and checked what packages were involved in the solution of &lt;code&gt;ocaml&lt;/code&gt; as this would represent a minimum set of packages that would affect virtually all packages. The list was big, but not &lt;em&gt;too&lt;/em&gt; big:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;winpthreads, system-msvc, system-mingw, ocaml-variants, ocaml-system,
ocaml-options-vanilla, ocaml-option-tsan, ocaml-option-static,
ocaml-option-spacetime, ocaml-option-no-flat-float-array,
ocaml-option-no-compression, ocaml-option-nnpchecker,
ocaml-option-nnp, ocaml-option-musl, ocaml-option-mingw,
ocaml-option-leak-sanitizer, ocaml-option-fp, ocaml-option-flambda,
ocaml-option-default-unsafe-string, ocaml-option-bytecode-only,
ocaml-option-afl, ocaml-option-address-sanitizer,ocaml-option-32bit,
ocaml-config, ocaml-compiler, ocaml-beta, ocaml-base-compiler, ocaml,
dkml-base-compiler, conf-unwind, conf-pkg-config, base-unix,
base-threads, base-ocamlbuild, base-nnp, base-metaocaml-ocamlfind,
base-implicits, base-effects, base-domains, base-bigarray
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I tried the same thing whilst using the oxcaml opam-repository, and this time, the list became much &lt;em&gt;much&lt;/em&gt; larger:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;zed, zarith-xen, zarith-freestanding, zarith, yojson, xenstore, xdg,
x509, webbrowser, wasm_of_ocaml-compiler, variantslib, uutf, uuseg,
uunf, uucp, uTop, uri-sexp, uri, uopt, univ_map, uchar, tyxml,
typerex, typerep, trie, topkg, tls-lwt, tls, timezone, time_now,
textutils_kernel, textutils, tcpip, system-msvc, system-mingw,
swhid_core, stringext, string_dict, stdune, stdlib-shims, stdio, ssl,
splittable_random, spdx_licenses, spawn, shell,
shared-memory-ring-lwt, shared-memory-ring, sha, sexplib0, sexplib,
sexp_pretty, seq, sedlex, rresult, result, regex_parser_intf,
record_builder, react, re2, re, randomconv, publish, ptime, psq,
protocol_version_header, ppxlib_jane, ppxlib_ast, ppxlib, ppxfind,
ppx_yojson_conv_lib, ppx_yojson_conv, ppx_variants_conv, ppx_var_name,
ppx_typerep_conv, ppx_typed_fields, ppx_tydi, ppx_tools_versioned,
ppx_tools, ppx_template, ppx_string_conv, ppx_string,
ppx_stable_witness, ppx_stable, ppx_shorthand, ppx_sexp_value,
ppx_sexp_message, ppx_sexp_conv, ppx_pipebang, ppx_optional,
ppx_optcomp, ppx_module_timer, ppx_log, ppx_let, ppx_js_style,
ppx_jane, ppx_inline_test, ppx_ignore_instrumentation, ppx_here,
ppx_helpers, ppx_hash, ppx_globalize, ppx_fixed_literal,
ppx_fields_conv, ppx_fail, ppx_expect, ppx_enumerate,
ppx_disable_unused_warnings, ppx_diff, ppx_deriving, ppx_derivers,
ppx_custom_printf, ppx_cstruct, ppx_compare, ppx_cold, ppx_bin_prot,
ppx_bench, ppx_base, ppx_assert, pp, portable, pipe_with_writer_error,
pcre, pbkdf, patch, parsexp, ounit2, ordering, optint, opam-state,
opam-repository, opam-publish, opam-lib, opam-format,
opam-file-format, opam-core, ojs, ohex, odoc-parser, odoc, octavius,
ocplib-endian, ocp-indent, ocp-build, ocb-stubblr, ocamlnet,
ocamlgraph, ocamlformat-rpc-lib, ocamlformat-lib, ocamlformat,
ocamlfind-secondary, ocamlfind, ocamlc-loc, ocamlbuild,
ocaml_intrinsics_kernel, ocaml_intrinsics, ocaml-version,
ocaml-variants, ocaml-system, ocaml-syntax-shims,
ocaml-secondary-compiler, ocaml-options-vanilla, ocaml-option-tsan,
ocaml-option-static, ocaml-option-spacetime,
ocaml-option-no-flat-float-array, ocaml-option-no-compression,
ocaml-option-nnpchecker, ocaml-option-nnp, ocaml-option-musl,
ocaml-option-mingw, ocaml-option-leak-sanitizer, ocaml-option-fp,
ocaml-option-flambda, ocaml-option-default-unsafe-string,
ocaml-option-bytecode-only, ocaml-option-afl,
ocaml-option-address-sanitizer, ocaml-option-32bit,
ocaml-migrate-parsetree, ocaml-lsp-server, ocaml-index,
ocaml-freestanding, ocaml-config, ocaml-compiler-libs,
ocaml-base-compiler, ocaml, obuild, num, nocrypto, mtime, mmap,
mirage-xen-posix, mirage-xen, mirage-types, mirage-time, mirage-stack,
mirage-solo5, mirage-sleep, mirage-runtime, mirage-random,
mirage-ptime, mirage-protocols, mirage-profile, mirage-no-xen,
mirage-no-solo5, mirage-net-xen, mirage-net, mirage-mtime,
mirage-kv-mem, mirage-kv-lwt, mirage-kv, mirage-flow, mirage-entropy,
mirage-device, mirage-crypto-rng-mirage, mirage-crypto-rng-lwt,
mirage-crypto-rng, mirage-crypto-pk, mirage-crypto-ec, mirage-crypto,
mirage-clock-unix, mirage-clock-lwt, mirage-clock, mew_vi, mew,
metrics-lwt, metrics, merlin-lib, merlin, menhirSdk, menhirLib,
menhirCST, menhir, mdx, magic-mime, macaddr-cstruct, macaddr, lwt_ssl,
lwt_react, lwt_ppx, lwt_log, lwt-dllist, lwt, lsp, lru, logs,
lambda-term, kdf, jst-config, jsonrpc, jsonm, js_of_ocaml-toplevel,
js_of_ocaml-ppx, js_of_ocaml-lwt, js_of_ocaml-compiler, js_of_ocaml,
jbuilder, jane_rope, jane-street-headers, ipaddr-sexp, ipaddr-cstruct,
ipaddr, io-page, int_repr, http, hkdf, hex, hacl_x25519, gmap,
github-unix, github-data, github, gen_js_api, gen, gel,
functoria-runtime, fpath, fmt, fix, fieldslib, fiber, fiat-p256,
ezjsonm, extlib-compat, extlib, expectree, expect_test_helpers_core,
ethernet, eqaf, either, easy-format, dyn, duration, dune-site,
dune-rpc, dune-release, dune-private-libs, dune-configurator,
dune-compiledb, dune-build-info, dune, dot-merlin-reader, domain-name,
dkml-base-compiler, digestif, curly, cstruct-sexp, cstruct-lwt,
cstruct, csexp, crunch, cpuid, cppo, core_unix, core_kernel,
core_extended, core, configurator, conf-which, conf-unwind,
conf-pkg-config, conf-ninja, conf-m4, conf-libssl, conf-libpcre,
conf-gmp-powm-sec, conf-gmp, conf-g++, conf-cmake, conf-c++,
conf-bash, conf-autoconf, conduit-lwt-unix, conduit-lwt, conduit,
cohttp-lwt-unix, cohttp-lwt-jsoo, cohttp-lwt, cohttp, cmdliner,
cmarkit, chrome-trace, charInfo_width, capitalization, camomile,
camlp4, camlp-streams, ca-certs, bos, biniou, binaryen-bin, bin_prot,
bigstringaf, bigarray-compat, bheap, basement, base_quickcheck,
base_bigstring, base64, base-unix, base-threads, base-ocamlbuild,
base-num, base-nnp, base-effects, base-domains, base-bytes,
base-bigarray, base, backoff, atdgen-runtime, atdgen, atd, async_unix,
async_rpc_kernel, async_log, async_kernel, async_extra, async,
astring, asn1-combinators, arp, angstrom, alcotest
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This enormous list is because the opam file for oxcaml - &lt;code&gt;ocaml-variants.5.2.0+ox&lt;/code&gt; - lists a bunch of conflicts to ensure that various incompatible packages are never selected:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;conflicts: [
  &amp;quot;base&amp;quot; {&amp;lt; &amp;quot;v0.18~&amp;quot;}
  &amp;quot;alcotest&amp;quot; {!= &amp;quot;1.9.0+ox&amp;quot;}
  &amp;quot;backoff&amp;quot; {!= &amp;quot;0.1.1+ox&amp;quot;}
  &amp;quot;dot-merlin-reader&amp;quot; {!= &amp;quot;5.2.1-502+ox&amp;quot;}
  &amp;quot;gen_js_api&amp;quot; {!= &amp;quot;1.1.2+ox&amp;quot;}
  &amp;quot;js_of_ocaml&amp;quot; {!= &amp;quot;6.0.1+ox&amp;quot;}
  &amp;quot;js_of_ocaml-compiler&amp;quot; {!= &amp;quot;6.0.1+ox&amp;quot;}
  &amp;quot;js_of_ocaml-ppx&amp;quot; {!= &amp;quot;6.0.1+ox&amp;quot;}
  &amp;quot;js_of_ocaml-toplevel&amp;quot; {!= &amp;quot;6.0.1+ox&amp;quot;}
  &amp;quot;jsonrpc&amp;quot; {!= &amp;quot;1.19.0+ox&amp;quot;}
  &amp;quot;lsp&amp;quot; {!= &amp;quot;1.19.0+ox&amp;quot;}
  &amp;quot;lwt_ppx&amp;quot; {!= &amp;quot;5.9.1+ox&amp;quot;}
  &amp;quot;mdx&amp;quot; {!= &amp;quot;2.5.0+ox&amp;quot;}
  &amp;quot;merlin&amp;quot; {!= &amp;quot;5.2.1-502+ox&amp;quot;}
  &amp;quot;merlin-lib&amp;quot; {!= &amp;quot;5.2.1-502+ox&amp;quot;}
  &amp;quot;ocaml-compiler-libs&amp;quot; {!= &amp;quot;v0.17.0+ox&amp;quot;}
  &amp;quot;ocaml-index&amp;quot; {!= &amp;quot;1.1+ox&amp;quot;}
  &amp;quot;ocaml-lsp-server&amp;quot; {!= &amp;quot;1.19.0+ox&amp;quot;}
  &amp;quot;ocamlbuild&amp;quot; {!= &amp;quot;0.15.0+ox&amp;quot;}
  &amp;quot;ocamlformat&amp;quot; {!= &amp;quot;0.26.2+ox&amp;quot;}
  &amp;quot;ocamlformat-lib&amp;quot; {!= &amp;quot;0.26.2+ox&amp;quot;}
  &amp;quot;ojs&amp;quot; {!= &amp;quot;1.1.2+ox&amp;quot;}
  &amp;quot;ppxlib&amp;quot; {!= &amp;quot;0.33.0+ox&amp;quot;}
  &amp;quot;ppxlib_ast&amp;quot; {!= &amp;quot;0.33.0+ox&amp;quot;}
  &amp;quot;sedlex&amp;quot; {!= &amp;quot;3.3+ox&amp;quot;}
  &amp;quot;topkg&amp;quot; {!= &amp;quot;1.0.8+ox&amp;quot;}
  &amp;quot;uTop&amp;quot; {!= &amp;quot;2.15.0+ox&amp;quot;}
  &amp;quot;uutf&amp;quot; {!= &amp;quot;1.0.3+ox&amp;quot;}
  &amp;quot;wasm_of_ocaml-compiler&amp;quot; {!= &amp;quot;6.0.1+ox&amp;quot;}
  &amp;quot;zarith&amp;quot; {!= &amp;quot;1.12+ox&amp;quot;}
]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and it seems that the solver is looking not just at these packages, but also at all of their dependencies too. So this is a much larger set of packages that we need to track changes for, probably making the caching an awful lot less effective. It&apos;s not clear to me that this is the best way for the solver to handle conflicts, but I don&apos;t know enough about how it works yet to say for sure.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/09/caching-opam-solutions.html" rel="alternate" title="Caching opam solutions"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/09/build-ids-for-day10.html</id>
    <title type="text">Build IDs for Day10</title>
    <updated>2025-09-08T00:00:00Z</updated>
    <published>2025-09-08T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;&lt;a href=&quot;https://tunbury.org&quot;&gt;mtelvers&lt;/a&gt;, &lt;a href=&quot;https://www.dra27.uk/blog/&quot;&gt;dra27&lt;/a&gt; and I have been working on a system to build opam packages similar to the way that the docs-ci system does - effectively building a per-package binary cache to do very fast builds of the entire opam repository. It supports building even mutually-incompatible packages by dynamically creating the build environment for each package, and thus allows us to generate something akin to &lt;a href=&quot;https://github.com/ocurrent/opam-health-check&quot;&gt;opam health check&lt;/a&gt; but much faster.&lt;/p&gt;
&lt;p&gt;Currently the cache of a package is a key-value store where the key is a hash of the package name and version and all of its dependencies and their name and version, alongside some information about the OS. This is great when this info can uniquely identify the output, but this isn&apos;t always the case. In particular, the oxcaml opam-repository has several packages where the version number is the upstream version number with `-ox` appended, as they have patches to make them compatible with oxcaml. If these patches change without bumping the suffix the currently caching mechanism would lead to trouble. When we discussed this David pointed out the idea of the &lt;a href=&quot;https://github.com/ocaml/opam/blob/c36dd1ce40a715ef27122184715bbf3e9aa7f0c9/src/state/opamPackageVar.ml#L178-L211&quot;&gt;build-id&lt;/a&gt; in opam, which would perfectly satisfy our needs. Unfortunately this code is quite deep within the opam codebase and at the point we need it we don&apos;t have an installed opam switch, so we need to pull the code out and insert it into our project.&lt;/p&gt;
&lt;p&gt;One of the first challenges was that day10 currently includes the OS details in the hash so that we can test across different distros. This is at odds with the opam build-id which doesn&apos;t include that, so in order to try to get as close as possible to the opam hash I split the cache into 2 layers - a per-OS cache directory containing hashes based on pure opam metadata. The idea is that these should be identical to the build-ids of opam. With that fixed, the new cache layout looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;debian-12-x86_64/123...abc/{build.log,config,...}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where the &lt;code&gt;123...abc&lt;/code&gt; should be the same as the build-id you would get with all the packages contained installed.&lt;/p&gt;
&lt;p&gt;Now my actual use case for this is to track the state of the oxcaml world day by day, so for this I need to track both the opam-repository for OCaml and also the opam repository for OxCaml. The project currently uses a Makefile for coordinating the builds, but I thought it was time we moved on to a dedicated batch execution process. So I asked Claude to knock me up one of those, using odoc_driver for inspiration. It&apos;s very basic right now, simply iterating through the latest versions of every package, but I have got it to check on cache hits and misses, so I should be able to run it tomorrow to see how quickly we can test PRs to oxcaml/opam-repository&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/09/build-ids-for-day10.html" rel="alternate" title="Build IDs for Day10"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/09/giving-hub-cl-an-upgrade.html</id>
    <title type="text">Giving hub.cl an upgrade</title>
    <updated>2025-09-07T00:00:00Z</updated>
    <published>2025-09-07T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;For a few years now we&apos;ve been running &lt;code&gt;hub.cl.cam.ac.uk&lt;/code&gt;, a Jupyterhub instance, for the first year course &amp;quot;Foundations of Computer Science&amp;quot;. It serves as a hosting site for the lecture notes, which come in the form of Jupyter notebooks, and as a playground where students can try OCaml, and it also is used to run the assessed exercises that are a mandatory part of the course.&lt;/p&gt;
&lt;p&gt;Since I spent some time setting it up back in 2018 or so, its aggregated some cruft over the years, and has also fallen somewhat behind the bleeding edge of the Jupyter software stack. So I thought this year, as I&apos;m actually lecturing the course, I&apos;d give it a bit of loving care and attention.&lt;/p&gt;
&lt;p&gt;We were still on Jupyterhub 1.5.3 whereas the current release is 5.3.0 - so there was quite a bit of work to do. I brief play with putting things on the latest version seemed to break quite a lot of things, so I thought it might be better to go back to the drawing board and start the config again from scratch. So with some help from Claude, I&apos;ve now managed to hugely simplify the whole config of Jupyterhub, and even given it a makeover to try to match the style of www.cst.cam.ac.uk as well. The improvements include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Using caddy as a reverse proxy for TLS termination, meaning I don&apos;t have to manually renew the letsencrypt cert every 3 months&lt;/li&gt;
&lt;li&gt;Unifying the configuration of the two container images used for students and instructors&lt;/li&gt;
&lt;li&gt;Upgrading to much newer jupyterhub, notebook and nbgrader images&lt;/li&gt;
&lt;li&gt;Simplifying the configuration required to make it work on a new server - persistent user directories are now docker volumes rather than bindmounts on the local filesystem&lt;/li&gt;
&lt;li&gt;Updating the authentication method to use Raven via OAuth2 rather than the unmaintained &lt;a href=&quot;https://github.com/pyCav/jupyterhub-raven-auth&quot;&gt;jupyterhub-raven-auth&lt;/a&gt; which I&apos;d had to maintain &lt;a href=&quot;https://github.com/jonludlam/jupyterhub-raven-auth/commit/36eaf16b410e7ac3cfc532269e0ae5f1de34f231&quot;&gt;a patch&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Rebasing &lt;a href=&quot;https://github.com/jonludlam/nbgrader/commit/c83a6cbb7b530ce87b0b157accddcdc832bcba38&quot;&gt;my patch&lt;/a&gt; to nbgrader to verify all of the output of the cells when grading answers
As ever, this took longer than I&apos;d anticipated, but I&apos;m mostly there now. There are a few more steps to try:&lt;/li&gt;
&lt;li&gt;trial the &lt;a href=&quot;https://github.com/akabe/ocaml-jupyter/pull/210&quot;&gt;new patch&lt;/a&gt; for using ocaml-jupyter with OCaml 5.x&lt;/li&gt;
&lt;li&gt;see how to upgrade to notebook v7, as I&apos;ve stuck with v6 in order to keep the extensions we&apos;re using going.&lt;/li&gt;
&lt;/ul&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/09/giving-hub-cl-an-upgrade.html" rel="alternate" title="Giving hub.cl an upgrade"/>
    <category term="teaching"/>
    <category term="ocaml"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/08/ocaml-lsp-mcp.html</id>
    <title type="text">Using ocaml-lsp-server via an MCP server</title>
    <updated>2025-08-27T00:00:00Z</updated>
    <published>2025-08-27T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Here&apos;s a quick post on how to get the OCaml Language Server (ocaml-lsp-server) working with an MCP server.&lt;/p&gt;
&lt;p&gt;We&apos;re going to use &lt;a href=&quot;https://github.com/isaacphi&quot;&gt;issacphi&lt;/a&gt;&apos;s adapter for LSP servers, which is written in go. So install go, and then:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;go install github.com/isaacphi/mcp-language-server@latest
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Once that&apos;s done, make sure you&apos;ve got `ocaml-lsp-server` installed in your switch:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;opam install ocaml-lsp-server
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then add the MCP config for claude where you want to run it:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;claude mcp add ocamllsp -s local -t stdio -- /Users/jon/go/bin/mcp-language-server -workspace . -lsp ocamllsp
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It&apos;d be nice to get this working `globally` - that is, with `-s user` - but I haven&apos;t been able to get that to work yet.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/08/ocaml-lsp-mcp.html" rel="alternate" title="Using ocaml-lsp-server via an MCP server"/>
    <category term="ai"/>
    <category term="plugins"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/08/ocaml-mcp-server.html</id>
    <title type="text">An OCaml MCP server</title>
    <updated>2025-08-20T00:00:00Z</updated>
    <published>2025-08-20T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;LLMs are proving themselves superbly capable of a variety of coding tasks, having been trained against the enormous amount of code, tutorials and manuals available online. However, with smaller languages like OCaml there simply isn&apos;t enough training material out there, particularly when it comes to new language features like &lt;a href=&quot;https://ocaml.org/manual/5.3/effects.html&quot;&gt;effects&lt;/a&gt; or new packages that haven&apos;t had time to be widely used. With my colleagues &lt;a href=&quot;https://anil.recoil.org/&quot;&gt;Anil&lt;/a&gt;, &lt;a href=&quot;https://ryan.freumh.org/&quot;&gt;Ryan&lt;/a&gt; and &lt;a href=&quot;https://toao.com/&quot;&gt;Sadiq&lt;/a&gt; we&apos;ve been exploring ways to &lt;a href=&quot;https://anil.recoil.org/notes/cresting-the-ocaml-ai-hump&quot;&gt;improve this situation&lt;/a&gt;. One way we can mitigate these challenges is to provide a Model Context Protocol (&lt;a href=&quot;https://modelcontextprotocol.io&quot;&gt;MCP&lt;/a&gt;) server that&apos;s capable of providing up-to-date info on the current state of the OCaml world.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;https://docs.anthropic.com/en/docs/mcp&quot;&gt;MCP specification&lt;/a&gt; was released by Anthropic at the end of last year. Since then it has become an astonishingly popular mechanism for extending the capabilities of LLMs, allowing them to become incredibly powerful agents capable of much more than simply chatting. There are now a huge variety of MCP servers, from one that provides &lt;a href=&quot;https://github.com/r-huijts/firstcycling-mcp&quot;&gt;professional cycling data&lt;/a&gt; to one that can &lt;a href=&quot;https://github.com/GongRzhe/Gmail-MCP-Server&quot;&gt;do your email&lt;/a&gt;. The &lt;a href=&quot;https://github.com/punkpeye/awesome-mcp-servers&quot;&gt;awesome mcp server list&lt;/a&gt; already lists hundreds, and these are just the &lt;em&gt;awesome&lt;/em&gt; ones!&lt;/p&gt;
&lt;p&gt;I&apos;ve been working with &lt;a href=&quot;https://toao.com/&quot;&gt;Sadiq&lt;/a&gt; to make an &lt;a href=&quot;https://github.com/sadiqj/odoc-llm/&quot;&gt;MCP server for OCaml&lt;/a&gt;, with an initial focus on building it such that it can be hosted for everyone rather than something that is run locally. Our plan is to start with a service that can help with choosing OCaml libraries, by taking advantage of the work done by &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci/&quot;&gt;ocaml-docs-ci&lt;/a&gt; which is the tool used to generate the documentation for all packages in &lt;a href=&quot;https://github.com/ocaml/opam-repository&quot;&gt;opam-repository&lt;/a&gt; and is served by &lt;a href=&quot;https://ocaml.org/&quot;&gt;ocaml.org&lt;/a&gt;. As well as producing HTML docs, we can also extract a number of other formats from the pipeline, including a newly created &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1341&quot;&gt;markdown backend&lt;/a&gt;. Using this, we can get markdown-formatted documentation for the every version of every package in the OCaml ecosystem.&lt;/p&gt;
&lt;h2&gt;Semantic searching&lt;/h2&gt;
&lt;p&gt;The first thing we focused on was being able to do a &lt;em&gt;semantic search&lt;/em&gt; over the whole OCaml ecosystem. To do this, we&apos;re using &lt;a href=&quot;https://huggingface.co/spaces/hesamation/primer-llm-embedding&quot;&gt;LLM embeddings&lt;/a&gt;, for which we need some natural-language description to seach through.&lt;/p&gt;
&lt;p&gt;The documentation produced by &lt;code&gt;ocaml-docs-ci&lt;/code&gt; is generated per library module using &lt;a href=&quot;https://github.com/ocaml/odoc&quot;&gt;odoc&lt;/a&gt;, relying on the package author to provide documentation comments for each element in the signature. However, even if the package authors &lt;em&gt;hasn&apos;t&lt;/em&gt; provided any documentation, we can still see the types, values, modules and so on that the library exposes, and this is often enough to get a good idea of what the module does. We then take these documentation pages, which are formatted in markdown, and summarise them via an LLM at the module level. This is done hierarchically, so we start with the &apos;deepest&apos; modules, and then insert their summaries into the text of their parent module, then summarise those and so on. We found it useful to include the names and &lt;a href=&quot;https://ocaml.github.io/odoc/odoc/odoc_for_authors.html#preamble&quot;&gt;preambles&lt;/a&gt; of the ancestor modules when doing the summarisation to give additional context to the LLM. For example, here is the prompt generated for a submodule of the &lt;a href=&quot;https://erratique.ch/software/astring&quot;&gt;astring&lt;/a&gt; library:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-markdown&quot;&gt;Module: Astring.String.Ascii

Ancestor Module Context:
- Astring: Alternative `Char` and `String` modules. Open the module to
use it. This defines one value in your scope, redefines the `(^)`
operator, the `Char` module and the `String` module. Consult the
differences with the OCaml `String` module, the porting guide and a
few examples.
- Astring.String: Strings, `substrings`, string sets and maps. A
string `s` of length `l` is a zero-based indexed sequence of `l`
bytes. An index `i` of `s` is an integer in the range [`0`;`l-1`], it
represents the `i`th byte of `s` which can be accessed using the
string indexing operator `s.[i]`.
Important. OCaml&apos;s `string`s became immutable since 4.02. Whenever
possible compile your code with the `-safe-string` option. This module
does not expose any mutable operation on strings and assumes strings
are immutable. See the porting guide.

Module Documentation: US-ASCII string support.
References.

## Predicates
- val is_valid : string -&amp;gt; bool (* `is_valid s` is `true` iff only for
  all indices `i` of `s`, `s.[i]` is an US-ASCII character, i.e. a
  byte in the range [`0x00`;`0x7F`]. *)

## Casing transforms
The following functions act only on US-ASCII code points that is on
bytes in range [`0x00`;`0x7F`], leaving any other byte intact. The
functions can be safely used on UTF-8 encoded strings; they will of
course only deal with US-ASCII casings.

- val uppercase : string -&amp;gt; string (* `uppercase s` is `s` with
  US-ASCII characters `&apos;a&apos;` to `&apos;z&apos;` mapped to `&apos;A&apos;` to `&apos;Z&apos;`. *)
- val lowercase : string -&amp;gt; string (* `lowercase s` is `s` with
  US-ASCII characters `&apos;A&apos;` to `&apos;Z&apos;` mapped to `&apos;a&apos;` to `&apos;z&apos;`. *)
- val capitalize : string -&amp;gt; string (* `capitalize s` is like
  `uppercase` but performs the map only on `s.[0]`. *)
- val uncapitalize : string -&amp;gt; string (* `uncapitalize s` is like
  `lowercase` but performs the map only on `s.[0]`. *)

## Escaping to printable US-ASCII
- val escape : string -&amp;gt; string (* `escape s` is `s` with: *)
- val unescape : string -&amp;gt; string option (* `unescape s` unescapes
  what `escape` did. The letters of hex escapes can be upper, lower or
  mixed case, and any two letter hex escape is decoded to its
  corresponding byte. Any other escape not defined by `escape` or
  truncated escape makes the function return `None`. *)
- val escape_string : string -&amp;gt; string (* `escape_string s` is like
  `escape` except it escapes `s` according to OCaml&apos;s lexical
  conventions for strings with: *)
- val unescape_string : string -&amp;gt; string option (* `unescape_string`
  is to `escape_string` what `unescape` is to `escape` and also
  additionally unescapes the sequence `&amp;quot;\\&apos;&amp;quot;` (`0x5C,0x27`) to `&amp;quot;&apos;&amp;quot;`
  (`0x27`). *)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where clearly the package author has provided excellent documentation comments. This is then passed to an LLM which generated the following description:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;This module provides functions to check if a string contains only
US-ASCII characters, convert case for ASCII letters, and escape or
unescape strings using ASCII conventions. It operates on standard
OCaml strings, treating them as sequences of bytes, and ensures
compatibility with UTF-8 encoded strings when transforming case. Use
cases include sanitizing input for ASCII-only protocols, preparing
strings for environments requiring strict ASCII formatting, and
handling escaped string representations in configuration or
serialization contexts.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Once we have these natural language descriptions, we can generate embeddings for them to allow for semantic search amongst all modules in opam.&lt;/p&gt;
&lt;p&gt;In addition to the module descriptions, we also generate similar natural-language descriptions of the &lt;em&gt;package&lt;/em&gt; as a whole, by taking the README from the package and summarising it similarly. Where there is no README, we summarise the summaries of the modules of the libraries, so we&apos;re always able to generate some text description of the entire package.&lt;/p&gt;
&lt;p&gt;To help with the ranking, we&apos;re also using a measure of popularity for both modules and packages. For packages, we&apos;re using the number of reverse dependencies in opam as a proxy for popularity, and for modules, we&apos;re using the &amp;quot;occurrences&amp;quot; generated as part of the docs build. These [occurrences] are a count of how often modules are used in other modules, and are calculated by looking at the compiled [cmt] files and resolving references to external modules using odoc&apos;s internal logic and counting them.&lt;/p&gt;
&lt;p&gt;Once we have both the module and package summaries, we generate an embedding of the descriptions to allow for a semantic search to be performed efficiently. We&apos;re using this in two ways - to search for packages for broad queries of functionality, which just uses the package summaries, and for more specific queries to search for modules within packages.&lt;/p&gt;
&lt;p&gt;For the module search, if the packages to search in haven&apos;t been specified, we search for both modules and packages and then combine the results. This is particularly helpful when the search is for generic functionality that might be found in more specific packages. For example, a module-only search for the term &amp;quot;time and date manipulation functions&amp;quot; returns the strongest match with a &lt;a href=&quot;https://ocaml.org/p/caqti/2.2.4/doc/caqti.platform/Caqti_platform/Conv/index.html&quot;&gt;module from caqti&lt;/a&gt;, which, as caqti is a library for talking to relational databases, might not be what the user is looking for.&lt;/p&gt;
&lt;p&gt;We then put these search tools into an MCP server, along with a little more functionality. The server currently provides these five functions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Search for OCaml packages&lt;/li&gt;
&lt;li&gt;Search for OCaml modules (optionally within packages)&lt;/li&gt;
&lt;li&gt;Get the summary description of a package&lt;/li&gt;
&lt;li&gt;Get the raw Markdown docs for a module (optionally within a package)&lt;/li&gt;
&lt;li&gt;Search using sherldoc
The first 2 use the LLM-generated summaries as described above, and the last is using &lt;a href=&quot;https://github.com/art-w/&quot;&gt;Arthur&apos;s&lt;/a&gt; &lt;a href=&quot;https://github.com/art-w/sherlodoc&quot;&gt;sherlodoc tool&lt;/a&gt; which can do various searches, including type-based search, across the output of the &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci&quot;&gt;ocaml-docs-ci&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Example searches&lt;/h2&gt;
&lt;p&gt;The following are the results from some example package searches:&lt;/p&gt;
&lt;h3&gt;&amp;quot;HTTP client&amp;quot;&lt;/h3&gt;
&lt;pre&gt;&lt;code class=&quot;language-nolang&quot;&gt;#1 - http (v6.1.1)
  Similarity: 0.7593
  Reverse Dependencies: 407
  Combined Score: 0.6588
  Description: This package provides a comprehensive OCaml library for
building HTTP clients and servers with support for multiple
asynchronous programming model s. It enables developers to implement
efficient, portable HTTP services using different backends such as
Lwt, Async, Eio, and JavaScript, making it suitable for both Unix and
browser environments. The library emphasizes performance, modularity,
and interoperability, allowing custom backend implementations and
seamless in tegration with other OCaml libraries. It is commonly used
in web services, API clients, standalone microkernels, and
OCaml-to-JavaScript compilations for web app lications.

#2 - cohttp (v6.1.1)
  Similarity: 0.7377
  Reverse Dependencies: 403
  Combined Score: 0.6435
  Description: This package provides a comprehensive library for
 building HTTP clients and servers in OCaml. It supports multiple
asynchronous programming models and backends, enabling flexible
development across different runtime environments. The library offers
efficient handling of HTTP/1.1 and HTTPS, with portable pa rsing and
modular architecture. It is widely used for web services, API clients,
and standalone network applications.

#3 - cohttp-lwt-unix (v6.1.1)
  Similarity: 0.7089
  Reverse Dependencies: 338
  Combined Score: 0.6212
  Description: This package provides an implementation of the Cohttp
library using the Lwt asynchronous programming framework with Unix
bindings. It enables buil ding efficient HTTP clients and servers in
OCaml, supporting both synchronous and asynchronous network
operations. The package handles core HTTP functionality, i ncluding
request and response parsing, connection management, and HTTPS support
via OCaml-TLS. It is suitable for applications requiring
high-performance web ser vices, microservices, or networked
applications in the OCaml ecosystem.

#4 - cohttp-lwt (v6.1.1)
  Similarity: 0.7067
  Reverse Dependencies: 367
  Combined Score: 0.6207
  Description: This package provides a comprehensive library for
 building HTTP clients and servers in OCaml, supporting multiple
asynchronous programming models. It enables developers to implement
 efficient, portable HTTP services with support for both synchronous
 and asynchronous I/O, including secure HTTPS communicatio n. The
 package includes backends for Lwt, Async, Mirage, JavaScript, and
 Eio, making it versatile for use in different runtime environments,
 from Unix servers to web browsers. It is well-suited for applications
 requiring high-performance networking, such as web services, API
 clients, and embedded networked systems.

#5 - quests (v0.1.3)
  Similarity: 0.7960
  Reverse Dependencies: 1
  Combined Score: 0.6180
  Description: This package provides a high-level HTTP client library
for making web requests in OCaml. It simplifies interacting with HTTP
 servers by offering a n intuitive API for common methods like GET and
 POST, supporting features such as query parameters, form and JSON
 data submission, and automatic handling of gzip compression and
 redirects. It also includes authentication mechanisms like basic and
 bearer tokens, with partial support for sessions. Typical use cases
 include consuming REST APIs, scraping web content, or integrating
 with web services securely and efficiently.

#6 - ezcurl (v0.2.4)
  Similarity: 0.7395
  Reverse Dependencies: 6
  Combined Score: 0.5979
  Description: This package provides a simplified interface for making
HTTP requests in OCaml, built on top of the OCurl library. It
addresses the need for an ea sy-to-use, reliable, and stable API for
handling common web interaction tasks, such as fetching URLs and
processing responses. The package supports both synchron ous and
asynchronous operations, enabling efficient handling of parallel
requests and non-blocking I/O. Practical use cases include web
scraping, API client deve lopment, and integrating HTTP-based services
into OCaml applications.


{2 &amp;quot;Cryptographic hash&amp;quot;}

{@nolang[
#1 - digestif (v1.3.0)
  Similarity: 0.8165
  Reverse Dependencies: 621
  Combined Score: 0.7041
  Description: This package provides a comprehensive implementation of
  cryptographic hash functions, supporting algorithms such as MD5,
  SHA1, SHA2, SHA3, WHIRLPOOL, BLAKE2, and RIPEMD160. It allows users
  to choose between C and OCaml backends at link time, offering
  flexibility in performance and deployment scenarios. The library is
  designed for applications requiring secure hashing, such as data
  integrity verification, digital signatures, and cryptographic
  protocols. It is well-suited for systems programming and
  security-related applications in the OCaml ecosystem.

#2 - ppx_hash (vv0.17.0)
  Similarity: 0.7284
  Reverse Dependencies: 3337
  Combined Score: 0.6833
  Description: This package generates efficient hash functions for
  OCaml types based on their structure, enabling precise control over
  hashing behavior. It addresses the limitations of OCaml&apos;s built-in
  polymorphic hashing by allowing users to define custom hash
  functions during type derivation. Key features include selective
  field ignoring, support for folding-style hash accumulation, and
  compatibility with comparison and serialization systems. It is
  suitable for use with hash tables, persistent data structures, and
  any application requiring deterministic, type-driven hashing.

#3 - ez_hash (v0.5.3)
  Similarity: 0.8366
  Reverse Dependencies: 3
  Combined Score: 0.6583
  Description: This package provides a straightforward interface to
  common cryptographic hash functions, simplifying their use in OCaml
  applications. It wraps secure, widely-used algorithms like SHA-256
  and Blake2b, offering consistent and safe APIs for hashing data. The
  library is designed for clarity and ease of integration, making it
  ideal for developers needing reliable cryptographic operations
  without deep expertise in security. Practical uses include data
  integrity verification, digital signatures, and secure data storage.

#4 - murmur3 (v0.3)
  Similarity: 0.7805
  Reverse Dependencies: 1
  Combined Score: 0.6072
  Description: This package provides OCaml bindings for MurmurHash, a
fast and widely used non-cryptographic hash function. It enables
 efficient hash value compu tation for arbitrary data, making it
suitable for applications like hash tables, checksums, and data
fingerprinting. The bindings offer consistent hashing across platforms
and integrate seamlessly into OCaml projects requiring
high-performance hashing. Use cases include caching, distributed
systems, and data integrity ve rification where cryptographic security
is not required.

#5 - kdf (v1.0.0)
  Similarity: 0.6775
  Reverse Dependencies: 473
  Combined Score: 0.6033
  Description: This package implements standard key derivation
functions (KDFs) for cryptographic applications in OCaml. It supports
scrypt, PBKDF1, PBKDF2, and HKDF, enabling secure generation of
cryptographic keys from passwords or shared secrets. These functions
help mitigate brute-force attacks and ensure keys are de rived in a
reproducible, secure manner. Use cases include password-based
encryption, secure token generation, and key material expansion in
cryptographic protocols.
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Module-level search : &amp;quot;time and date manipulation functions&amp;quot;&lt;/h3&gt;
&lt;pre&gt;&lt;code class=&quot;language-nolang&quot;&gt;#1 - timmy-jsoo: Timmy_jsoo
  Similarity: 0.5460
  Original Similarity: 0.7800
  Popularity Score: 0.0000
  Description: This module provides precise date and time arithmetic,
conversion, and comparison operations across multiple representations,
including OCaml-nati ve, JavaScript, and string formats. It works with
structured types like `Date.t`, `Time.t`, and ISO weeks, supporting
timezone-aware transformations and RFC3339 formatting. Concrete use
cases include cross-runtime timestamp synchronization, calendar-aware
scheduling, and robust temporal data validation in distributed
systems.

#2 - calendar: CalendarLib
  Similarity: 0.5331
  Original Similarity: 0.7616
  Popularity Score: 0.3448
  Description: This module provides precise date and time manipulation
  with support for calendar operations, time zones, periods, and
  formatted input/output. It works with types like `Calendar.t`,
  `Date.t`, `Time.t`, and `Period.t` to handle tasks such as event
  scheduling, timestamp conversion, and historical date calculations.
  Concrete use cases include scheduling systems, log timestamping,
  holiday calculations, and cross-timezone time normalization.

#3 - calendar: CalendarLib.Fcalendar
  Similarity: 0.5191
  Original Similarity: 0.6820
  Popularity Score: 0.1390
  Description: This module provides float-based calendar operations
  for date creation, conversion, and manipulation, including time zone
  adjustments, component extraction (year/month/day/hour/second), and
  arithmetic with periods. It works with a `t` type representing time
  as float seconds, alongside `day`, `month`, `year`, and Unix time
  structures, prioritizing Unix time precision over sub-second
  accuracy. It suits applications tolerating minor imprecision in date
  comparisons or arithmetic, such as logging systems or coarse-grained
  scheduling, where exact floating-point equality isn&apos;t critical.

#4 - calendar: CalendarLib.Calendar_builder.Make
  Similarity: 0.5112
  Original Similarity: 0.7302
  Popularity Score: 0.0785
  Description: This module combines date and time functionality to
  construct and manipulate calendar values with float-based precision,
  offering operations like timezone conversion, component extraction
  (day, month, year, etc.), and arithmetic using `Period.t`. It works
  with a calendar type `t` that integrates date and time components,
  alongside conversions to Unix timestamps, Julian day numbers, and
  structured representations like `Unix.tm`. Designed for scenarios
  requiring precise temporal calculations (e.g., calendar arithmetic,
  Gregorian date validation, or leap day checks), it balances
  flexibility with known precision limitations inherent to float-based
  time representations.

#5 - timmy-unix: Clock
  Similarity: 0.5080
  Original Similarity: 0.7257
  Popularity Score: 0.0000
  Description: This module provides functions to retrieve the current
  POSIX time, the local timezone, and the current date in the local
  timezone. It works with time and date types from the Timmy library,
  specifically `Timmy.Time.t` and `Timmy.Date.t`. Use this module to
  obtain precise time and date information for logging, scheduling, or
  time-based computations.
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Module-level search: &amp;quot;Balanced Tree&amp;quot;&lt;/h3&gt;
&lt;pre&gt;&lt;code class=&quot;language-nolang&quot;&gt;#1 - grenier: Mbt
  Similarity: 0.5274
  Original Similarity: 0.7534
  Popularity Score: 0.0495
  Description: This module implements a balanced binary tree structure
  with efficient concatenation and size-based operations. It supports
  tree construction through leaf and node functions, automatically
  balancing nodes and annotating them with values from a provided
  measure module. It is useful for applications requiring fast access,
  dynamic sequence management, and efficient merging of tree-based
  data structures.

#2 - camomile: CamomileLib.AvlTree
  Similarity: 0.5008
  Original Similarity: 0.7155
  Popularity Score: 0.0495
  Description: This module implements balanced binary trees (AVL
  trees) with operations for constructing, deconstructing, and
  traversing trees. It supports key operations like inserting nodes,
  extracting leftmost/rightmost elements, concatenating trees, and
  folding or iterating over elements. It is useful for maintaining
  ordered data with efficient lookup, insertion, and deletion, such as
  in symbol tables or priority queues.

#3 - batteries: BatAvlTree
  Similarity: 0.5003
  Original Similarity: 0.7147
  Popularity Score: 0.1485
  Description: This module implements balanced binary trees (AVL
  trees) with operations for creating, modifying, and traversing
  trees. It supports tree construction with optional rebalancing,
  splitting, and concatenation, and provides root, left, and right
  accessors with failure handling. Concrete use cases include
  efficient ordered key-value storage, set-like structures, and
  maintaining sorted data with logarithmic-time insertions and
  lookups.

#4 - grenier: Bt2
  Similarity: 0.4927
  Original Similarity: 0.7039
  Popularity Score: 0.2634
  Description: This module implements a balanced binary tree structure
  with efficient concatenation and rank-based access. It supports
  creating empty trees, constructing balanced nodes, and joining two
  trees with logarithmic cost relative to the smaller tree&apos;s size. Use
  cases include maintaining ordered collections with frequent splits
  and joins, and efficiently accessing elements by position.

#5 - grenier: Mbt.Make
  Similarity: 0.4913
  Original Similarity: 0.7019
  Popularity Score: 0.0495
  Description: This module implements a balanced tree structure with
  efficient concatenation and size-based operations. It supports
  construction of trees using leaf and node functions, where nodes are
  automatically balanced and annotated with measurable values from
  module M. The module enables efficient rank queries and joining of
  trees, with applications in managing dynamic sequences where fast
  access and concatenation are critical.
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Limitations and future work&lt;/h2&gt;
&lt;p&gt;We&apos;re aware that there are currently a number of limitations with what&apos;s been done so far, and there&apos;s a lot of exciting things that could quite easily be added!&lt;/p&gt;
&lt;p&gt;We haven&apos;t done much prompt optimisation either for the tools themselves, nor their descriptions in the MCP server. We also haven&apos;t done much optimisation of the information retrieval - and it&apos;s clear from some of the results shown above that there are improvements to be made in the ranking algorithms. Some obvious next steps would be to do some &lt;a href=&quot;https://arxiv.org/html/2406.12433v2&quot;&gt;re-ranking&lt;/a&gt; or some form of hybrid search.&lt;/p&gt;
&lt;p&gt;A particular challenge is that since this is based entirely off of the &lt;code&gt;ocaml-docs-ci&lt;/code&gt; build, it won&apos;t necessarily reflect the actual API your local build, as for OCaml, this &lt;a href=&quot;https://jon.recoil.org/blog/2025/04/semantic-versioning-is-hard.html&quot;&gt;can&apos;t be done&lt;/a&gt;. Thibaut Mattio is working on a &lt;a href=&quot;https://github.com/tmattio/ocaml-mcp&quot;&gt;local MCP server&lt;/a&gt; that would be perfectly positioned to do some of what we&apos;re doing, although we&apos;d need to have a good local docs build implemented in dune for this to work well.&lt;/p&gt;
&lt;p&gt;Also, there&apos;s plenty more data that we&apos;ve collected during the docs builds! We can show the implementations of functions, we can expose code samples, select different versions of packages and much more. While we&apos;ve concentrated on the search aspects, there&apos;s still a lot of low-hanging fruit that can be worked on.&lt;/p&gt;
&lt;p&gt;If you&apos;re interested in helping us out on this, the project lives &lt;a href=&quot;https://github.com/sadiqj/odoc-llm&quot;&gt;on github&lt;/a&gt; - come along and join us!&lt;/p&gt;
&lt;h2&gt;Using the server&lt;/h2&gt;
&lt;p&gt;If you&apos;d like to try it, we&apos;ve got a demo server running right now. It&apos;s hosted on dill.caelum.ci.dev here at the Computer Laboratory in the University of Cambridge. To enable it with Claude, try this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;claude mcp add -t sse ocaml http://dill.caelum.ci.dev:8000/sse
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Obviously this is pre-alpha quality software, and we might take it down with no notice, and it might not work as expected, and all of the other usual caveats. Let us know if it works, or doesn&apos;t, or if you&apos;ve got some suggestions for improvements!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/08/ocaml-mcp-server.html" rel="alternate" title="An OCaml MCP server"/>
    <category term="ai"/>
    <category term="plugins"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/08/week33.html</id>
    <title type="text">Week 33</title>
    <updated>2025-08-19T00:00:00Z</updated>
    <published>2025-08-19T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;More work this week on the OCaml MCP server. Sadiq and I met before I went away on holiday and discussed the next steps to &apos;park&apos; the work on the MCP server. The final steps are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Write a README&lt;/li&gt;
&lt;li&gt;Write and run a small script to fix a problem with module-type names&lt;/li&gt;
&lt;li&gt;Write up and publish a blog post
Not much, right? As always though, writing things up lead to a whole load more work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The first problem occurred when writing up how it parsed the input docs. It turned out that when converting the repo so that it took markdown formatted files (using a &lt;a href=&quot;https://github.com/jonludlam/odoc/tree/odoc-llm-markdown&quot;&gt;slightly tweaked&lt;/a&gt; version of &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1341&quot;&gt;davesnx&apos;s PR&lt;/a&gt;), Claude had decided that the way to do this was to first convert the markdown into HTML, and then use the HTML parser it had already built. Whilst tidying this up, Claude was remarkably keen to just use regexps to parse the markdown rather than using a pre-existing markdown library, so it took a little persuasion to get it into a state I was happy with.&lt;/p&gt;
&lt;p&gt;The second issue was that the script that form the bulk of the repo had been written at different times, and therefore Claude didn&apos;t really take into account any of the decisions it had made in one script when building the next. So most of the command-line arguments were slightly different, which made writing up a mini &apos;howto&apos; in the README quite a jarring experience.&lt;/p&gt;
&lt;p&gt;Thirdly, and most importantly, we had decided that we needed a few example searches to show how the system worked. We&apos;d already had a &lt;a href=&quot;/blog/2025/07/week28.html&quot;&gt;useful experience&lt;/a&gt; with this when Anil had tried to search for a &apos;time and date parsing and formatting&apos; library, so it shouldn&apos;t really have been a surprise that trying a few more examples showed some more interesting behaviour. Specifically, the searches I wanted to do were for an &amp;quot;HTTP client&amp;quot;, &amp;quot;JSON parser&amp;quot;, &amp;quot;Cryptographic Hash&amp;quot; and Anil&apos;s time-and-date query, and in actually trying these searches and critically examining the results, I had to go back and figure out why they weren&apos;t giving me the results I had expected.&lt;/p&gt;
&lt;p&gt;The first of these searches I had anticipated would be quite interesting, as this is a query that should show the OCaml ecosystem &lt;a href=&quot;https://discuss.ocaml.org/t/simple-modern-http-client-library/11239&quot;&gt;missing an obvious HTTP client&lt;/a&gt;. However, even with this in mind one of the top results was one of Cohttp&apos;s module types, &lt;a href=&quot;https://ocaml.org/p/cohttp/latest/doc/cohttp/Cohttp/Generic/Client/module-type-S/index.html&quot;&gt;Cohttp.Generic.Client.S&lt;/a&gt;. This, of course, isn&apos;t much use if you&apos;re looking for an HTTP client, as module-types aren&apos;t going to give you an implementation to actually use. So I decided that we&apos;d exclude module-types from the results. This turned out to be slightly more tricky than I anticipated as we&apos;d lost the distinction between modules and module types further back in the pipeline, so Claude had to do some plumbing to ensure we had this information at the point we were doing the search.&lt;/p&gt;
&lt;p&gt;The cryptographic hash search gave some plausible looking results, so I moved on to the JSON search. I was expecting to see &lt;a href=&quot;/reference/yojson/yojson/Yojson/index.html&quot;&gt;&lt;code&gt;Yojson&lt;/code&gt;&lt;/a&gt; somewhere near the top of the list as that&apos;s a very popular library. I was also expecting to see &lt;a href=&quot;/reference/jsonm/jsonm/Jsonm/index.html&quot;&gt;&lt;code&gt;Jsonm&lt;/code&gt;&lt;/a&gt; somewhere near the top - or at least I&apos;d like to be able to find it by searching for a &apos;streaming parser&apos; as that&apos;s one of its key strengths. However, searching for &amp;quot;JSON parser&amp;quot; yielded some less than brilliant answers. The top 5 results were for modules in the packages &lt;code&gt;yojson-five&lt;/code&gt;, &lt;code&gt;decoders-yojson&lt;/code&gt;, &lt;code&gt;decoders-jsonaf&lt;/code&gt;, &lt;code&gt;ocplib-json-typed-browser&lt;/code&gt; and &lt;code&gt;ppx_protocol_conv_jsonm&lt;/code&gt;. While all of these are clearly in the same realm as I was after, having &lt;code&gt;jsonm&lt;/code&gt; show up literally 99th in the list, and &lt;code&gt;yojson&lt;/code&gt; itself not in the top 100 wasn&apos;t a great result.&lt;/p&gt;
&lt;p&gt;Some investigation showed that yojson had a particularly bad showing because the description of the module &lt;a href=&quot;/reference/yojson/yojson/Yojson/Basic/index.html&quot;&gt;&lt;code&gt;Yojson.Basic&lt;/code&gt;&lt;/a&gt; was the empty string! This turned out to be because of some bad error-handling logic in the summariser script, which ended up turning some errors into a blank description. Since running the summariser costs actual money, I didn&apos;t want to just rerun the whole thing, so I asked Claude for a script to find these problems and rerun them. The problem is not totally trivial as the summaries of child modules are used when generating the summary for parents, so when one is regenerated we should regenerate the summaries of all ancestors too. Given my recent experiences with Claude I&apos;d like to look this over quite carefully before letting it loose on my data, so I&apos;ve run it on yojson, which seemed to do the right thing, but not yet on the rest of the packages.&lt;/p&gt;
&lt;p&gt;Having fixed this, I still found that &lt;code&gt;jsonm&lt;/code&gt; was making a very poor showing. This turned out to be because the description it gives itself is a &amp;quot;Non-blocking streaming JSON codec for OCaml&amp;quot; which had a fairly low similarity with &amp;quot;JSON parser&amp;quot;. I was using a fairly small embedding model for the queries - Qwen/Qwen3-Embedding-0.6B, so I thought I might address this by using a larger one, and opted for Qwen/Qwen3-Embedding-8B. The machine I had been using for the MCP server has no GPU and had taken a while to do the embeddings using the 0.6B model, so I switched to generating them on my M4 macbook. This went &lt;em&gt;much&lt;/em&gt; faster, though since I have about 70Mb of module summaries it still took quite a while. This improved the situation somewhat, but it was still not high in the list.&lt;/p&gt;
&lt;p&gt;So I took a step back and had a think about the problem some more. Searching for a JSON parser is really quite a high-level search, and when evaluating the results I realised I was really thinking in terms of packages rather than modules. So I thought we could split the search in two - a package search and a module search. The package search would be used for the broad queries where you&apos;re interested in pulling in whole chunks of functionality, and the module search is for more low-level queries. In fact, the &apos;time and dating formatting&apos; query is somewhere in between, so I might need to have some more example queries for the module search functions. In addition, the module search could be restricted to the set of packages you&apos;re using, which might make it even more useful.&lt;/p&gt;
&lt;p&gt;Part of the split meant that I needed a different source of &apos;popularity&apos; for the packages than the occurrences data that came out of docs ci, as that was per-module and I needed something per-package. The obvious thing is to look at reverse dependencies in opam. I have this kind-of working, but it&apos;s currently not particularly smart, so this will need a little more attention. For example, it currently thinks that &lt;a href=&quot;https://melange.re/v5.0.0/&quot;&gt;melange&lt;/a&gt; has over 3000 reverse dependencies.&lt;/p&gt;
&lt;p&gt;With these changes in place, a package search for &apos;JSON parser&apos; now returns &lt;code&gt;yojson&lt;/code&gt; as number one, followed by &lt;code&gt;ppx_deriving_yojson&lt;/code&gt;, &lt;code&gt;ezjsonm&lt;/code&gt;, &lt;code&gt;ocplib-json-typed&lt;/code&gt; and &lt;code&gt;jsonaf&lt;/code&gt;. Unfortunately &lt;code&gt;jsonm&lt;/code&gt; is still languishing in 27th place, so there&apos;s still some tweaking to do.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/08/week33.html" rel="alternate" title="Week 33"/>
    <category term="weeknotes"/>
    <category term="ai"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/07/retrospective.html</id>
    <title type="text">4 months in, a retrospective</title>
    <updated>2025-07-18T00:00:00Z</updated>
    <published>2025-07-18T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Astonishingly, it&apos;s already been &lt;em&gt;four whole months&lt;/em&gt; since starting back at the university, which I find incredibly hard to believe. I&apos;m utterly convinced that it was only a couple of weeks ago that I walked back into the Computer Laboratory as an SRA for the first time since 2021, but here we are, at the end of term already. Time to do a bit of a retrospective and forward-looking plan for the next 3-4 months!&lt;/p&gt;
&lt;h2&gt;What&apos;s happened?&lt;/h2&gt;
&lt;p&gt;On wednesday this week, I had a chance to sit down with Anil, supposedly to talk about the upcoming lecturing of 1A Foundations of Computer Science, but we ended up talking about what I&apos;ve been doing for the past few months, and where it fits into the broader picture of the group as a whole. It was a really useful conversation, and I thought it would be good to outline it here while it&apos;s fresh in my mind.&lt;/p&gt;
&lt;p&gt;So then, to start, what have I been doing? What have I achieved? What have I learnt? It&apos;s been a bit of a daunting experience, landing in a team that are already working one hundred miles an hour on things well out of my comfort zone. I&apos;ve been going to group meetings and having lots of interesting conversations, but I&apos;ve found it difficult to make the next steps happen. One area where I&apos;ve had some success is in working with Sadiq on LLMs - in particular, getting local LLMs to solve programming exercises that we both &lt;a href=&quot;https://toao.com/blog/ocaml-local-code-models&quot;&gt;wrote&lt;/a&gt; &lt;a href=&quot;/blog/2025/05/ticks-solved-by-ai.html&quot;&gt;up&lt;/a&gt;. I&apos;ve also been working with him on taking the output from the docs CI and &lt;a href=&quot;https://github.com/sadiqj/odoc-llm&quot;&gt;summarising it with LLMs&lt;/a&gt; in order to create an MCP server that would help tools like &lt;a href=&quot;https://anthropic.com/&quot;&gt;Claude Code&lt;/a&gt; to choose OCaml packages to solve users&apos; problems.&lt;/p&gt;
&lt;p&gt;It&apos;s been somewhat easier, partly due to inertia, to carry on with projects that had been in flight at the time I started. Things like getting the Odoc 3 generated docs onto ocaml.org, which is finally complete only &lt;a href=&quot;/blog/2025/07/odoc-3-live-on-ocaml-org.html&quot;&gt;as of this week!&lt;/a&gt;. This has taken a whole lot of time, but I&apos;m really pleased with the end results. There&apos;s still an awful lot of improvements that I&apos;d like to see made, which, after drawing breath for a couple of weeks, I&apos;ll be writing down.&lt;/p&gt;
&lt;p&gt;An itch I&apos;d been wanting to scratch for a long time has been to look at client-side ocaml notebooks. I decided to make this an integral &lt;a href=&quot;/blog/2025/04/this-site.html&quot;&gt;feature of this blog&lt;/a&gt;, and I&apos;ve learnt an awful lot doing it. An important feature of this that I&apos;ve been keeping in mind is the idea that we could use the ocaml-docs-ci tool to build the libraries, which would allow us to host a toplevel for every single package in opam-repository - allowing at best &lt;a href=&quot;https://discuss.ocaml.org/t/an-example-for-every-ocaml-package/16953/10&quot;&gt;interactive examples&lt;/a&gt;, and at bare minimum merlin for live type-checking and autocompletion. The important principles to keep in mind for this are that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We have one &apos;toplevel&apos; javascript file, and libraries and cmis are dynamically loaded&lt;/li&gt;
&lt;li&gt;The interface between the frontend and the worker must not rely on a matched pair, e.g. an OCaml-5.3-compiled frontend might be talking to an OCaml-4.08-compiled worker thread - or even an oxcaml one!
I have this all working on my blog, where I have both an oxcaml worker and a standard ocaml worker and they both dynamically load in libraries and cmis as specified on the page.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I&apos;ve also supervised a 1A course for the first time - &lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2425/IntroProb/&quot;&gt;Introduction to Probability&lt;/a&gt;, and I&apos;ve done some marking for the 1A Foundations of Computer Science.&lt;/p&gt;
&lt;p&gt;Something that I&apos;d been expecting to do a lot on was work with oxcaml, but as the release happened later than anticipated and it coinciding with the marking and supervising, I&apos;ve not done quite as much of this as I had thought I would. In addition, I had anticipated working on Odoc to start implementing the new features of oxcaml, but to avoid duplicating effort I&apos;ve been waiting for the patches that have already been written at Jane Street to at least get odoc to compile, which have taken longer than I had hoped to get to me.&lt;/p&gt;
&lt;h2&gt;What&apos;s next?&lt;/h2&gt;
&lt;p&gt;With that in mind, Anil and I then talked about the bigger picture, as those of you who know Anil will be entirely unsurprised to hear! In particular, how will we be weaving the various threads of these activites - the teaching of OCaml, the large-scale (for OCaml) CI work, the LLMs and Oxcaml work together to form a coherent whole? How do I find a balance between them and ensure that we find &lt;a href=&quot;https://arxiv.org/abs/1106.0848&quot;&gt;synergies&lt;/a&gt; as opposed to pulling in different directions? How do make sure what we&apos;re doing helps us navigate the upending of the nature of development that agentic coding is bringing?&lt;/p&gt;
&lt;h3&gt;Efficient and reusable CI&lt;/h3&gt;
&lt;p&gt;A clear and obvious area where we&apos;ll be able to see real progress is to extract from docs CI the logic that I&apos;ve been using to do efficient builds of packages. As I previously &lt;a href=&quot;/blog/2025/07/odoc-3-live-on-ocaml-org.html&quot;&gt;wrote about&lt;/a&gt;, the new CI system is far more efficient than some of the other ocurrent-based pipelines, and it would save a huge amount of compute time if we were to take this tech and apply it elsewhere.&lt;/p&gt;
&lt;p&gt;So, how might we take what we&apos;ve got and produce something useful to everyone? We need to take a hammer to the fracture points of the docs CI service and split it into individually useful parts. Here are some next steps as I see them now. Let&apos;s take the solver out of docs CI, and have a service whose sole job is to create a repository of up-to-date solutions for all versions of all packages in opam-repository. These are the data that allow us to build the tree of package builds.&lt;/p&gt;
&lt;p&gt;Next, turn these solutions into one giant build. Perhaps a script? Maybe a giant buildkit dockerfile? This is very similar to Mark Elvers&apos; &lt;a href=&quot;https://github.com/mtelvers/ohc&quot;&gt;day10&lt;/a&gt; project. We can get this running on a big machine and just see how fast we can build everything. The key thing here is that it should be &lt;em&gt;trivial&lt;/em&gt; to run this on a linux box. A raspberry pi or a 768-core behemoth with 3TiB of ram. Just how fast &lt;em&gt;can&lt;/em&gt; we get it going? It&apos;s already building in a couple of days using &lt;a href=&quot;/blog/2025/07/odoc-3-live-on-ocaml-org.html&quot;&gt;sage&lt;/a&gt;, but that&apos;s using ocurrent/obuilder, which isn&apos;t quite the right tool for the job, and on a relatively puny machine. Can we do it in an hour? 10 minutes? Certainly the incrememntal builds ought to be done in seconds. What&apos;s the limit?&lt;/p&gt;
&lt;p&gt;These tools can then be used as the foundation for other CI systems. For opam-repo-ci, where we should be able to do the builds for a new package incredibly quickly. For opam-health-check, where we currently build foundational packages like dune and findlib &lt;em&gt;thousands of times&lt;/em&gt; per run.&lt;/p&gt;
&lt;p&gt;Once we&apos;ve got the packages built, docs CI is simply a pass over the top of the built artifacts. ocaml-docs-ci already demonstrates this - it only takes a few hours to rebuild all the docs when a new version of odoc is released, but in a way that only benefits docs! All the CI systems should be able to use this.&lt;/p&gt;
&lt;p&gt;We should also then be able to run js_of_ocaml on the libraries to build to infrastructure needed for the per-package toplevels for ocaml.org that I mentioned above. Each of these steps should be separate stages in a pipeline - one where each step produces artifacts for the next to consume.&lt;/p&gt;
&lt;p&gt;When we mix in some of the projects that other people in the team are working on, like David&apos;s work on &lt;a href=&quot;https://www.dra27.uk/blog/&quot;&gt;relocatable OCaml&lt;/a&gt;, we&apos;ve got something that might be able to form a basis for a binary cache for Dune Package Management, particularly when we involve Ryan&apos;s &lt;a href=&quot;https://ryan.freumh.org/papers.html#2025-arxiv-hyperres&quot;&gt;Hyperres&lt;/a&gt; paper so we might check that dependencies from outside of the OCaml universe are correct. Can we use &lt;a href=&quot;https://github.com/quantifyearth/shark&quot;&gt;Patrick and Michael&apos;s shark&lt;/a&gt; to do the build steps? Can we use these images to serve up toplevels for ocaml.org that are &lt;em&gt;real toplevels&lt;/em&gt; rather than javascript toplevels? Can we use these build environments to do help with reinforcement learning to train LLMs on OCaml code? There are a lot of interesting directions to take this work.&lt;/p&gt;
&lt;h3&gt;Other projects&lt;/h3&gt;
&lt;p&gt;There are, of course, other responsibilities that I have. Some of these I&apos;ll be able to fit in with the theme above, and some - well - maybe I&apos;ll have to figure out how to delegate them, a skill that I am not particularly good at, but one that I feel I should learn!&lt;/p&gt;
&lt;h4&gt;Teaching&lt;/h4&gt;
&lt;p&gt;A looming, terrifying, but tremendously exciting opportunity is teaching of 1A Foundations of Computer Science. This is amongst the first courses we teach our incoming undergraduates, currently lectured by &lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2425/FoundsCS/&quot;&gt;Anil&lt;/a&gt;. As he&apos;s on sabbatical this year, he has asked me to step up and lecture it. This is definitely not one for delegation!&lt;/p&gt;
&lt;p&gt;The immediate question, partly raised by my work with Sadiq, is: what do we do about LLMs? How should we adjust our teaching to take into account the existence of these tools? We had a very interesting chat earlier in the term with Professor &lt;a href=&quot;https://eecs.iisc.ac.in/people/prof-viraj-kumar/&quot;&gt;Viraj Kumar&lt;/a&gt; from &lt;a href=&quot;https://eecs.iisc.ac.in/&quot;&gt;IISc&lt;/a&gt; who was visiting Cambridge earlier this year. He&apos;s been &lt;a href=&quot;https://dl.acm.org/doi/10.1145/3724363.3729100&quot;&gt;working on this question&lt;/a&gt; for a while now, and I hope to have some more conversations with him over the summer.&lt;/p&gt;
&lt;h4&gt;Odoc paper&lt;/h4&gt;
&lt;p&gt;An area where I&apos;ve really made a shockingly small amount of progress is to write up all the work that&apos;s gone into Odoc over the past 6 (!!!) years.&lt;/p&gt;
&lt;h4&gt;Odoc notebooks&lt;/h4&gt;
&lt;p&gt;This needs to be tidied up and a v0.1 released. In particular, the work on js_top_worker might well be shared with Arthur&apos;s &lt;a href=&quot;https://github.com/art-w/x-ocaml&quot;&gt;x-ocaml&lt;/a&gt; for a unified toplevel experience.&lt;/p&gt;
&lt;h4&gt;AI work&lt;/h4&gt;
&lt;p&gt;I&apos;d like to carry on the work I&apos;ve started with Sadiq on the interaction of LLMs with OCaml. Getting the package search to work sensibly for an MCP server is first on the list, but also doing some reinforcement learning to improve specifically the perfomance on OCaml is very interesting, but not something I&apos;ve managed to carve out the time for yet. Something along the lines of &lt;a href=&quot;https://arxiv.org/abs/2504.21798&quot;&gt;swesmith&lt;/a&gt; but adapted for OCaml.&lt;/p&gt;
&lt;h4&gt;Oxcaml Odoc&lt;/h4&gt;
&lt;p&gt;Odoc needs to have some work done on it to support the new work that&apos;s gone into oxcaml, for example documenting of the modes. This is something I do expect to be working on soon.&lt;/p&gt;
&lt;h4&gt;Dune and odoc&lt;/h4&gt;
&lt;p&gt;Work needs to be done on the dune rules for odoc, which currently only support the feature-set in odoc 2.x. Paul-Elliot has &lt;a href=&quot;https://github.com/ocaml/dune/pull/11716&quot;&gt;done some work on this&lt;/a&gt;, but much more needs to be done.&lt;/p&gt;
&lt;h4&gt;Further general odoc work&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Better source rendering&lt;/li&gt;
&lt;li&gt;Syntax for linking to source&lt;/li&gt;
&lt;li&gt;Custom tags (used in odoc_notebook)&lt;/li&gt;
&lt;li&gt;Web-native rendering, for embedding odoc in a website&lt;/li&gt;
&lt;li&gt;Unifying paths and cpaths (https://github.com/jonludlam/odoc/tree/parameterised-paths)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;What to &lt;em&gt;actually&lt;/em&gt; do?&lt;/h2&gt;
&lt;p&gt;There are a lot of things in the above list. I&apos;m not sure yet how I manage to figure out what I actually end up doing, and how that helps me to help Tarides, to fit in as a useful member of the EEG group, and to make sure I&apos;m doing what&apos;s right for my own future. I feel the core project of the CI work will help everyone, but slotting the other work into the bigger picture will require some careful thought.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/07/retrospective.html" rel="alternate" title="4 months in, a retrospective"/>
    <category term="meta"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/07/odoc-3-live-on-ocaml-org.html</id>
    <title type="text">Odoc 3 is live on OCaml.org!</title>
    <updated>2025-07-14T00:00:00Z</updated>
    <published>2025-07-14T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;As of today, Odoc 3 is now live on OCaml.org! This is a major update to odoc, and has brought a whole host of new features and improvements to the documentation pages.&lt;/p&gt;
&lt;p&gt;Some of the highlights include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Source code rendering&lt;/li&gt;
&lt;li&gt;Hierarchical manual pages&lt;/li&gt;
&lt;li&gt;Image, video and audio support&lt;/li&gt;
&lt;li&gt;Separation of API docs by library
A huge amount of work went into the &lt;a href=&quot;https://discuss.ocaml.org/t/ann-odoc-3-0-released/16339&quot;&gt;Odoc 3.0 release&lt;/a&gt;, and I&apos;d like to thank my colleagues at Tarides, in particular &lt;a href=&quot;https://github.com/panglesd&quot;&gt;Paul-Elliot&lt;/a&gt; and &lt;a href=&quot;https://github.com/julow/&quot;&gt;Jules&lt;/a&gt; for the work they put into this.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But the odoc release happened months ago, so why is it only going live now? So, the doc tool itself is only one small part of getting the docs onto ocaml.org. Odoc works on the &lt;a href=&quot;https://discuss.ocaml.org/t/cmt-cmti-question/5308&quot;&gt;cmt and cmti&lt;/a&gt; files that are produced during the build process, and so part of the process of building docs is to build the packages, so we have to, at minimum, attempt to build all 17,000 or so distinct versions of the packages in opam-repository. The &lt;a href=&quot;https://github.com/ocurrent&quot;&gt;ocurrent&lt;/a&gt; tool &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci&quot;&gt;ocaml-docs-ci&lt;/a&gt;, which I&apos;ve previously &lt;a href=&quot;/blog/2025/05/docs-progress.html&quot;&gt;written&lt;/a&gt; &lt;a href=&quot;/blog/2025/04/ocaml-docs-ci-and-odoc-3.html&quot;&gt;about&lt;/a&gt;, is responsible for these builds and in this new release has demonstrated a new approach to this task, where we attempt to do the build in as efficient a way as possible by effectively building binary packages once for each required package in a specific &apos;universe&apos; of dependencies. For example, many packages require e.g. &lt;a href=&quot;https://erratique.ch/software/cmdliner&quot;&gt;cmdliner.1.3.0&lt;/a&gt; to build, and some require a specific version of OCaml too. So we&apos;ll build cmdliner.1.3.0 once against each version of OCaml required -- but &lt;em&gt;only once&lt;/em&gt;, which is in contrast to how some of the other tools in the ocurrent suite work, e.g. &lt;a href=&quot;https://github.com/ocurrent/opam-repo-ci&quot;&gt;opam-repo-ci&lt;/a&gt;. Once the packages are built, we then run the new tool &lt;a href=&quot;https://ocaml.github.io/odoc/odoc-driver/index.html&quot;&gt;odoc_driver&lt;/a&gt; to actually build the HTML docs. In addition to this, a new feature of Odoc 3 is to be able to link to packages that are your direct dependencies - so for example, the docs of odoc contain links to the docs of odoc_driver, even though odoc_driver depends upon odoc. This, whilst sounding easy enough, required some radical changes in the docs ci, which I promise I will write about later!&lt;/p&gt;
&lt;p&gt;The builds and the generation of the docs is all done on a single blade server, called &lt;a href=&quot;https://sage.caelum.ci.dev&quot;&gt;sage&lt;/a&gt; with 40 threads, 2 8TiB spinning drives and a 1.8TiB SSD cache, and it produces about 1 TiB of data over the course of a couple of days. The changes required to this part of the process since odoc 2.x were primarily myself and &lt;a href=&quot;https://tunbury.org&quot;&gt;Mark Elvers&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Once the docs are built, how do they get onto ocaml.org? Odoc itself knows nothing about the layout and styling of ocaml.org, so the HTML it produces isn&apos;t suitable to be just rendered when a user requests particular docs. What happens is that odoc produces, as well as a self-contained HTML file, a json file with the body of the page, the sidebars, the breadcrumbs and so on as structured data, one per HTML page, which are then served from sage over HTTP. When a user requests a particular docs page, the ocaml.org server will request that json file from sage, then render this with the ocaml.org styling, then send it back to the user.&lt;/p&gt;
&lt;p&gt;As odoc 3 moved a fair bit of logic from ocaml.org into odoc itself, there were quite a few changes that needed to be made to the ocaml.org server to integrate this into the site. This work was mostly done by &lt;a href=&quot;https://github.com/panglesd&quot;&gt;Paul-Elliot&lt;/a&gt; and myself, with a lot of help from the &lt;a href=&quot;https://github.com/ocaml/ocaml.org?tab=readme-ov-file#maintainers&quot;&gt;ocaml.org team&lt;/a&gt;, in particular &lt;a href=&quot;https://github.com/sabine&quot;&gt;Sabine Schmaltz&lt;/a&gt; and &lt;a href=&quot;https://github.com/cuihtlauac&quot;&gt;Cuihtlauac Alvarado&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So, quite a lot of integration and infrastructure work was required to get the new docs site up and running, and I&apos;m very happy to see this particular task concluded!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/07/odoc-3-live-on-ocaml-org.html" rel="alternate" title="Odoc 3 is live on OCaml.org!"/>
    <category term="odoc"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/07/week28.html</id>
    <title type="text">Week 28</title>
    <updated>2025-07-14T00:00:00Z</updated>
    <published>2025-07-14T00:00:00Z</published>
    <content type="html">
      &lt;h2&gt;OCaml MCP server&lt;/h2&gt;
&lt;p&gt;Last week I got the summarisation to the point where it felt useful to run it across all the modules in opam. With this completed we then got to try out the MCP server to see how useful it would be in practice.&lt;/p&gt;
&lt;p&gt;One of the first queries &lt;a href=&quot;https://anil.recoil.org/&quot;&gt;Anil&lt;/a&gt; tried was to ask it which libraries would be most useful for &amp;quot;date time parsing and formatting&amp;quot;. We were surprised to see that the first two libraries it returned were &lt;code&gt;caqti&lt;/code&gt; and &lt;code&gt;mariadb&lt;/code&gt;, specifically mentioning the module &lt;a href=&quot;https://ocaml.org/p/caqti/latest/doc/caqti-platform/Caqti_platform/Conv/index.html&quot;&gt;Caqti_platform.Conv&lt;/a&gt; and &lt;a href=&quot;https://ocaml.org/p/mariadb/latest/doc/mariadb/Mariadb/module-type-S/Time/index.html&quot;&gt;Mariadb.S.Time&lt;/a&gt;. While these do indeed provide the required functionality, they&apos;re probably not the right libraries to provide this. It&apos;s going to be tricky to decide this in the MCP server, so we should probably be leaving it up to the LLM to decide amongst them on the client. However, for very general queries we might end up with a large number of matching libraries, so we&apos;ll need to have a limit on the number of packages returned, which implies some form of ranking.&lt;/p&gt;
&lt;p&gt;One way we can do this is by using the occurrences code in odoc. The idea is that we examine module implementation files (ie, ml rather than mli files), and counts the number of times the code uses values, types and other identifiers from other libraries. We can then aggregate these counts over all packages in opam repository and use it as an effective marker of popularity, which allows us to rank the results by popularity and only return the top N results.&lt;/p&gt;
&lt;p&gt;We&apos;re not currently using the occurrences for anything, so I wasn&apos;t especially surprised to find that it&apos;s not working as intended. There were a number of issues:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The occurrences output file was being written at a path not within the package dir, so it wasn&apos;t being persisted.&lt;/li&gt;
&lt;li&gt;The CLI interface for generating occurrences works by providing a directory containing the odocl files, but we were only providing the top-level directory and it wasn&apos;t recursively searching.&lt;/li&gt;
&lt;li&gt;Once the occurrences were captured, the aggregation step used the full identifier of the value being aggregated, meaning that, for example, &lt;a href=&quot;/reference/ocaml-compiler/stdlib/Stdlib/List/index.html#val-length&quot;&gt;&lt;code&gt;List.length&lt;/code&gt;&lt;/a&gt; in OCaml 5.3 was counted separately from &lt;a href=&quot;/reference/ocaml-compiler/stdlib/Stdlib/List/index.html#val-length&quot;&gt;&lt;code&gt;List.length&lt;/code&gt;&lt;/a&gt; in OCaml 4.14.
All of these issues are with code in the odoc repository, which, as it happens, also needs a release soon to ensure that it works with the imminent launch of OCaml 5.4. During the week, before I discovered the problems above, I had attempted to make a release of Odoc 3.1, but there was a license kerfuffle that, when combined with the issues in the occurrences code, gave me enough cause to pull the release.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Before I try to make the release again, this time I&apos;ll be running the release candidate with docs-ci, and checking that the occurrences make sense. I set this running on Friday afternoon, and it had completed by Friday evening, so it&apos;s actually pretty quick to rerun odoc on the 15,000 or so packages required for ocaml.org.&lt;/p&gt;
&lt;h2&gt;Trouble with this blog&lt;/h2&gt;
&lt;p&gt;In other news, in trying to post my blog at the beginning of the week, I was stymied a little by the changes in oxcaml. I had been using a custom opam-repository forked from the official oxcaml one, because I needed a patched js_of_ocaml in order to fix the toplevel code. I had hoped this would mean that I could update it on my schedule, rather than being at the mercy of upstream changes. Unfortunately though, the download URL for ocaml-flambda wasn&apos;t pointing at an immutable commit, so when I tried it I got a checksum error. So I ended up trying to rebase the changes onto the latest oxcaml opam-repository, which didn&apos;t go well at all. The version numbers had all changed, which in opam means that files are in different directories, so git got thoroughly confused. On top of that, because the js_of_ocaml repository has multiple packages in it, whereas opam repository has a directory per-package, we end up having multiple copies of the patches. So in the end I&apos;ve just committed all the patches to a git repo on github, and pinned it in the Dockerfile that builds this site.&lt;/p&gt;
&lt;p&gt;What would be handy is a way to apply the patches in a package in opam repository to and from a git repository, similar to quilt/guilt. We don&apos;t quite have all of the pieces to do this, as although we have a download URL and often a dev-repo, I don&apos;t believe we currently have a way to get the base commit of that repository.&lt;/p&gt;
&lt;h2&gt;Oxcaml continues&lt;/h2&gt;
&lt;p&gt;We had a meeting on Thursday with Jane Street on the next steps for oxcaml. There are a number of areas in which JS are keen for us to help out with.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Playgrounds - both javascript and docker-based. The playground on the oxcaml website right now uses github codespaces, which works nicely but currently takes an absolute age to start up. We can almost certainly improve this by building docker images and pushing them to the docker hub, rather than building oxcaml from scratch when starting the codespace. There&apos;s also interest in the javascript playgrounds, which can serve a slightly different purpose than the docker-based one, more limited in how it can be used, but without requiring someone to spin up a full docker container.&lt;/li&gt;
&lt;li&gt;Documentation - Odoc has had some patches to run on oxcaml, but there&apos;s no support for documenting many of the new features yet, including modes. We&apos;ve got to do some experiments here to see what the best way is to show the new type-system features in the generated docs. There were some suggestions of using javascript to show/hide the modes, for example.&lt;/li&gt;
&lt;li&gt;Improvements in Merlin - again this is an area ripe for investigation. In particular, how do we best expose the new features of the type system for users? What&apos;s needed here is user feedback from people who are actually using oxcaml to build real projects.&lt;/li&gt;
&lt;li&gt;Better error messages - OCaml has been getting improved error messages with each release, but there&apos;s still room for improvement, and the new features of the type system in particular have many different failure modes. Again, we need user feedback to understand the pain points and improve the error messages accordingly.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Next week&lt;/h2&gt;
&lt;p&gt;Next week, the plan is to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Check the occurrences from docs-ci, and integrate into the MCP server&lt;/li&gt;
&lt;li&gt;Talk to &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark&lt;/a&gt; about building the docker image for the oxcaml playground&lt;/li&gt;
&lt;li&gt;Tidy up the &lt;a href=&quot;/reference/js_top_worker/js_top_worker/Js_top_worker/index.html&quot;&gt;&lt;code&gt;Js_top_worker&lt;/code&gt;&lt;/a&gt; code so it can be used in the javascript oxcaml playground&lt;/li&gt;
&lt;li&gt;Release Odoc 3.1&lt;/li&gt;
&lt;/ul&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/07/week28.html" rel="alternate" title="Week 28"/>
    <category term="weeknotes"/>
    <category term="ai"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/07/week27.html</id>
    <title type="text">Weeks 24-27</title>
    <updated>2025-07-07T00:00:00Z</updated>
    <published>2025-07-07T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;It&apos;s been a busy few weeks. There&apos;s been exam marking for the 1A Foundations of Computer Science course, an Odoc release to plan, and some really interesting new work on using LLMs to summarise OCaml documentation. This post is about anaspect of that last one that I found particularly interesting.&lt;/p&gt;
&lt;h2&gt;odoc-llm&lt;/h2&gt;
&lt;p&gt;Sadiq and I have been &lt;a href=&quot;https://toao.com/blog/ocaml-local-code-models&quot;&gt;looking at using LLMs&lt;/a&gt; to summarise the documentation produced by Odoc. The idea is to get a concise summary of the purpose of each module, so that we can quickly identify which modules are relevant to a particular task.&lt;/p&gt;
&lt;p&gt;For testing this, we need to see how it works on different types of libraries. The first axis I wanted to test on goes between &apos;well documented&apos; and &apos;poorly documented&apos;, and so I need at least two libraries on opposite ends of the spectrum.&lt;/p&gt;
&lt;p&gt;For the &apos;well documented&apos; case, I chose &lt;code&gt;cmdliner&lt;/code&gt;. It&apos;s a library that I almost always have to look at the docs for when I&apos;m using it, as, despite using it many many times, the interface doesn&apos;t seem to stick in my head.&lt;/p&gt;
&lt;p&gt;For the &apos;poorly documented&apos; case, I chose &lt;code&gt;odoc&lt;/code&gt; itself, somewhat ironically. In defence of the odoc authors (me included!), the libraries it provides are simply there for code organisation and aren&apos;t meant to be consumed outside of the tool itself!&lt;/p&gt;
&lt;h3&gt;The approach&lt;/h3&gt;
&lt;p&gt;The output from Odoc is a set of HTML files, one per module/module type/class/etc., containing the documentation for that item. Our first take on this was to parse the HTML files and extract the text content, which we then fed to an LLM to summarise. However, this was pretty awkward, so we decided to try a PR that &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1341&quot;&gt;davesnx recently made to Odoc&lt;/a&gt; to output markdown instead of HTML.&lt;/p&gt;
&lt;p&gt;We look for leaf modules that don&apos;t contain any submodules, and start by summarising those, then move onto the parent modules, splicing in the summaries of the children, and so on, up to the top-level modules. We then move on to summarising the whole library, which usually is just a single namespace module, but occasionally is a group of top-level modules.&lt;/p&gt;
&lt;p&gt;One of the early prompts for the module &lt;code&gt;Cmdliner.Term.Syntax&lt;/code&gt; looked roughly as follows:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;You are an expert OCaml developer. Write a 2-3 sentence description focusing on:
- The specific operations and functions this module provides
- What data types or structures it works with
- Concrete use cases (avoid generic terms like &amp;quot;utility functions&amp;quot; or &amp;quot;common operations&amp;quot;)

Do NOT:
- Repeat the module name in the description
- Mention &amp;quot;functional programming patterns&amp;quot; or &amp;quot;code clarity&amp;quot;
- Use filler phrases like &amp;quot;provides functionality for&amp;quot; or &amp;quot;collection of functions&amp;quot;
- Describe how it works with other modules

Module: Syntax
Module Documentation: let operators.
( let+ ) is map.
( and* ) is product.
- val (let+) : &apos;a t -&amp;gt; (&apos;a -&amp;gt; &apos;b) -&amp;gt; &apos;b t (* ( let+ ) is map. *)
- val (and+) : &apos;a t -&amp;gt;
  &apos;b t -&amp;gt;
  (&apos;a * &apos;b) t (* ( and* ) is product. *)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and the output using a small model (qwen3-30b-a3b) was:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;&amp;quot;The module provides (let+) for applying functions to values within a context and (and+) for combining two contexts into a product. It operates on applicative
structures like option, list, or custom types that support these operations. For example, it enables sequential transformation of values in a context or
pairing elements from two separate contexts.&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;There are quite a few issues with the input here. Firstly, we&apos;ve only given the module name, not the full path. Secondly, there&apos;s nothing to let the model know what &lt;code&gt;t&lt;/code&gt; might be, so it has decided it&apos;s a completely generic type. It also has no idea about the context in which this module was found, so it has no idea that it&apos;s part of a command-line processing library. By fixing these issues, we end up with the prompt:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;You are an expert OCaml developer. Write a 2-3 sentence description focusing on:
- The specific operations and functions this module provides
- What data types or structures it works with
- Concrete use cases (avoid generic terms like &amp;quot;utility functions&amp;quot; or &amp;quot;common operations&amp;quot;)

Do NOT:
- Repeat the module name in the description
- Mention &amp;quot;functional programming patterns&amp;quot; or &amp;quot;code clarity&amp;quot;
- Use filler phrases like &amp;quot;provides functionality for&amp;quot; or &amp;quot;collection of functions&amp;quot;
- Describe how it works with other modules

Module: Cmdliner.Term.Syntax

Ancestor Module Context:
- Cmdliner: Declarative definition of command line interfaces.
Consult the tutorial, details about the supported command line syntax and examples of use.
Open the module to use it, it defines only three modules in your scope.
- Cmdliner.Term: Terms.
A term is evaluated by a program to produce a result, which can be turned into an exit status. A term made of terms referring to command line arguments implicitly defines a command line syntax.

Module Documentation: let operators.
- val (let+) : &apos;a Cmdliner.Term.t -&amp;gt; (&apos;a -&amp;gt; &apos;b) -&amp;gt; &apos;b Cmdliner.Term.t (* ( let+ ) is map. *)
- val (and+) : &apos;a Cmdliner.Term.t -&amp;gt;
  &apos;b Cmdliner.Term.t -&amp;gt;
  (&apos;a * &apos;b) Cmdliner.Term.t (* ( and* ) is product. *)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The output of this improved prompt is much better:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;The module provides operators to map and combine terms, which represent command line argument parsers and their results. `let+` transforms a parsed argument
into a new value, while `and+` merges two independent arguments into a tuple. These enable building structured command line interfaces, such as parsing a
filename and a flag simultaneously, then combining them into a configuration record.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It still occasionally incorrectly decides that this module provides monadic combinators rather than applicative, but this is where we get better results from using a larger model.&lt;/p&gt;
&lt;p&gt;There are quite a few other issues that we&apos;ve fixed - for example, treating module types differently than modules, and a bug where infix operators were being omitted from the documentation. In one case, it uncovered a bug in the markdown generator where includes weren&apos;t getting expanded, which I got &lt;a href=&quot;https://github.com/jonludlam/odoc/commit/926cca100c307818e57281c3d40e98f1975f0f95&quot;&gt;Claude to fix&lt;/a&gt;. My &lt;em&gt;modus operandi&lt;/em&gt; has essentially been to look at the output for the test packages, find a summary that looks bonkers, and then look back at the prompt to find that, indeed, the input was missing some crucial information.&lt;/p&gt;
&lt;p&gt;One thing I&apos;d quite like to try is to re-open a &lt;a href=&quot;https://github.com/ocaml/odoc/pull/655&quot;&gt;PR that Drup made&lt;/a&gt; as an April Fool&apos;s joke back in 2001, which ended up outputting OCaml formatted code. This is actually pretty close to what we might want to give to the LLM - though we&apos;d probably format the comments as markdown, and we&apos;d be replacing types with fully-qualified types as above. Funny how things turn out!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/07/week27.html" rel="alternate" title="Weeks 24-27"/>
    <category term="weeknotes"/>
    <category term="odoc"/>
    <category term="ai"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/06/week23.html</id>
    <title type="text">Week 23</title>
    <updated>2025-06-09T00:00:00Z</updated>
    <published>2025-06-09T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;merlinonly
Some brief notes on last week.&lt;/p&gt;
&lt;h2&gt;Docs CI and Sherlodoc&lt;/h2&gt;
&lt;p&gt;Anil has been working on an &lt;a href=&quot;https://tangled.sh/@anil.recoil.org/odoc-mcp&quot;&gt;MCP server&lt;/a&gt; that searches through the output of the docs CI to find relevant packages and API information for opam packages. For expediency, this works by scraping the HTML output, but a potentially better solution would be to integrate properly with &lt;a href=&quot;https://doc.sherlocode.com&quot;&gt;Sherlodoc&lt;/a&gt;, &lt;a href=&quot;https://github.com/art-w/&quot;&gt;Arthur&apos;s&lt;/a&gt; code search engine.&lt;/p&gt;
&lt;h3&gt;Building the index&lt;/h3&gt;
&lt;p&gt;To make this work with the new docs CI, first we need to build a sherlodoc search database. This involves grabbing all of the &lt;code&gt;.odocl&lt;/code&gt; files that odoc produces for the latest version of each library, copying them locally and running &lt;code&gt;sherlodoc index&lt;/code&gt; on the output. Getting &lt;em&gt;all&lt;/em&gt; of the odocl files is simple, but filtering so we only have the latest version is slightly more complex, as we need to use &lt;code&gt;opam&lt;/code&gt;&apos;s library to make sure we&apos;re looking at the latest versions.&lt;/p&gt;
&lt;p&gt;The simple way to get the odocl files ends up unpacking them into the filesystem in a directory hierarchy that matches the URL on ocaml.org, so we see files like:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;p/odoc/3.0.0/doc/odoc.document/odoc_document.odocl
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So finding the version number is a matter of listing the directories, for which I took &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci/blob/4dfe7e6265610da4e0ce2a386cfbf0b8eac3d9bd/src/lib/track.ml#L58-L76&quot;&gt;some code&lt;/a&gt; from docs CI:&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;type p = {
opam : OpamPackage.t;
path : Fpath.t;
}&lt;/p&gt;
&lt;p&gt;let rec take n l =
match n, l with
| n, x::xs when n &amp;gt; 0 -&amp;gt;
x :: take (n - 1) xs
| _, _ -&amp;gt; []&lt;/p&gt;
&lt;p&gt;let get_versions ~limit path =
let open Rresult in
let package = Fpath.basename path in
let mk_pkg v =
Printf.sprintf &amp;quot;%s.%s&amp;quot; package v
in
Bos.OS.Dir.contents path
&amp;gt;&amp;gt;| (fun versions -&amp;gt;
versions
|&amp;gt; List.map (fun path -&amp;gt;
{ opam = Fpath.basename path |&amp;gt; mk_pkg |&amp;gt; OpamPackage.of_string;
path = path })
)
|&amp;gt; Result.get_ok
|&amp;gt; (fun v -&amp;gt;
v
|&amp;gt; List.sort (fun a b -&amp;gt;
-OpamPackage.compare a.opam b.opam)
|&amp;gt; take limit)&lt;/x-ocaml&gt;
This gives us a sorted list of the versions for the package, and we can pick the first one to get the latest version. With the output from this we can then run &lt;code&gt;sherlodoc index&lt;/code&gt; and we get a nice big (1.7 gig!) index file.&lt;/p&gt;
&lt;h3&gt;Serving the index&lt;/h3&gt;
&lt;p&gt;The next step is to serve this index file so that the MCP server can access it. The file format is a marshalled OCaml value, so we need to have an OCaml program to read it in and perform the search, and it&apos;ll have to be a server since the whole index needs to be unmarshalled into memory before any search can be performed, and it would be dumb to do this for every query.&lt;/p&gt;
&lt;p&gt;Sherlodoc got partially integrated into odoc&apos;s code base before the 3.0 release with the exception of the server, which was left out to avoid pulling a load of new dependencies to odoc. Unfortunately, we didn&apos;t expose the sherlodoc libraries publicly, so we&apos;ll need to do that in order to make anything useful with sherlodoc. In addition, odoc embeds the version of odoc used into the odocl files, and since right now the docs CI is building with a version of odoc that &lt;em&gt;doesn&apos;t&lt;/em&gt; expose the libraries, we might have to hack around that in order to use those odocl files. Obviously the longer term solution is just to make a new release of odoc with this change and update the docs CI to use that.&lt;/p&gt;
&lt;h2&gt;Package to Library map&lt;/h2&gt;
&lt;p&gt;A related quest was to assemble a map of opam package to ocamlfind library names. It&apos;s a quirk of history that the library names that an opam package provides are not necessarily related to the name of the package. That means that tools like &lt;code&gt;dune&lt;/code&gt; have a hard time linting projects to check that the libraries they&apos;re using are mentioned in the opam files. Dune, of course, has resolved this be ensuring that it&apos;s an error to build a package using dune where the library names don&apos;t start with the package name, but as dune is just one of many OCaml build systems, the problem remains.&lt;/p&gt;
&lt;p&gt;Since docs CI has built every version of every package, and because the Odoc 3 package layout includes the library names in the paths to the documentation, we should be able to produce a fairly definitive list of the libraries that each package provides, which tools like dune can then use for this sort of lint check. We can just tweak the code above slightly to get the library names and output a big JSON file with the mapping - or perhaps with the exceptions to dune&apos;s rule.&lt;/p&gt;
&lt;p&gt;I thought this would be a neat first project to try Claude Code on - a &apos;starter for ten&apos; - as it were, so I signed up to use Claude code and gave it a shot.&lt;/p&gt;
&lt;p&gt;It handily produced a working program that did exactly what I wanted, including creating a test directory that it used to verify the code worked. One fascinating thing I noted as it scrolled past was that it tried to use &lt;code&gt;yojson&lt;/code&gt; to write the output, but failed to get it to work and reverted back to &lt;code&gt;printf&lt;/code&gt; output. I suspect this will be due to it finding it troublesome to figure out the various steps that need to be taken to use a new library in a dune project, so this is something to have a play with later.&lt;/p&gt;
&lt;p&gt;After a couple of iterations with different heuristics to disambiguate between library names and other directories, I got a working program producing a JSON file with only the exceptions to dune&apos;s rule. I took a look through and almost immediately found &lt;code&gt;camlidl&lt;/code&gt; suggesting it produces a library called &lt;code&gt;com&lt;/code&gt;. This didn&apos;t look right at all so I installed it and found that the library is actually named &lt;code&gt;camlidl&lt;/code&gt;. The &lt;code&gt;cma&lt;/code&gt; file, though, is named &lt;code&gt;com.cma&lt;/code&gt;, so it looks like &lt;code&gt;odoc_driver&lt;/code&gt; has a bug. Interestingly, running &lt;code&gt;odoc_driver&lt;/code&gt; locally gets the library name correct, so it&apos;s only an issue when running it in the docs CI. &lt;a href=&quot;https://github.com/ocaml/odoc/issues/1351&quot;&gt;Issue filed&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Further claude code experiments&lt;/h2&gt;
&lt;p&gt;To see how well Claude Code could handle more complex tasks, I thought I&apos;d give it a whirl on something more like its home territory, and somewhere where I was less familiar. I decided to ask it to write some code to make a nicer editor experience for the notebooks project. Since the &lt;a href=&quot;https://github.com/jonludlam/jsoo-code-mirror&quot;&gt;bindings to codemirror&lt;/a&gt; I&apos;m using are very minimal, any new features I want to use end up with needing to write a bunch of bindings first, and only then seeing if the feature works as I&apos;d like. So instead I thought I&apos;d get claude to write the editor code for me in javascript, and then I could make sure it works as I want and only then convert it to OCaml. This worked pretty nicely, and I&apos;ve now got a neat &lt;a href=&quot;https://jon.ludl.am/experiments/notebook-editor/notebook-editable.html&quot;&gt;demonstration editor&lt;/a&gt; that I can use to guide the OCaml implementation.&lt;/p&gt;
&lt;h2&gt;More notebook work&lt;/h2&gt;
&lt;p&gt;The oxcaml project will be launching this week hopefully. I&apos;ve been looking at Luke&apos;s &lt;a href=&quot;https://github.com/ocaml-flambda/flambda-backend/pull/3886&quot;&gt;Parallelism tutorial&lt;/a&gt; and have been thinking about how this will work with the online notebooks. The parallel library works via effects, and the oxcaml branch of js_of_ocaml doesn&apos;t support effects yet, and it might be a while before it does. However, the blog post is mainly talking about the intricacies of the type system work that&apos;s been done to ensure the parallel library is safe, and as such perhaps we can get a lot out of doing this online with just Merlin.&lt;/p&gt;
&lt;p&gt;Some early experimentation showed that the parallel library breaks the worker on load, so we need to do something a bit more sophisticated than just &apos;not call exec&apos;, so I did some work to have a mode of worker that doesn&apos;t load the cmas, just the cmis for Merlin. This is almost there.&lt;/p&gt;
&lt;h2&gt;Odoc work&lt;/h2&gt;
&lt;p&gt;Ocaml 5.4 is just around the corner, and there&apos;s some odoc work to be done before the release. One of the main new features that will impact odoc is the new &lt;a href=&quot;https://tyconmismatch.com/papers/ml2024_labeled_tuples.pdf&quot;&gt;labelled tuples&lt;/a&gt; feature. Fortunately &lt;a href=&quot;https://github.com/lukemaurer&quot;&gt;Luke Maurer&lt;/a&gt; has already done a lot of work to plumb this into odoc, so this will save us a lot of work - thanks, Luke! There&apos;s likely to be a few other bits and pieces to do, but hopefully not too much.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/06/week23.html" rel="alternate" title="Week 23"/>
    <category term="weeknotes"/>
    <category term="odoc"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/05/docs-progress.html</id>
    <title type="text">Progress in OCaml docs</title>
    <updated>2025-05-29T00:00:00Z</updated>
    <published>2025-05-29T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;The docs build is progress well, and we&apos;ve &lt;em&gt;just about&lt;/em&gt; hit 20,000 packages (20,038 to be precise). So at this point I thought it&apos;d be useful to take a look through the various failures to see if there are any insights to be gained.&lt;/p&gt;
&lt;p&gt;Odoc requires a built package in order to generate the docs, there are two steps that have to be done before we can begin building the docs. Step one is to figure out the exact set of packages to build - ie, doing an opam solve, and step two is to actually build the packages. These two steps are, to some extent, out of docs-ci&apos;s control, and rely on the state of opam repository. While there are efforts to keep this in as good a state as possible, it&apos;s still the case that these steps fail much more often than the actual docs build itself. Let&apos;s take a look at some of the failures we see in each of these steps.&lt;/p&gt;
&lt;h2&gt;Step 1: opam solve&lt;/h2&gt;
&lt;p&gt;There are 2,074 solver failures. A good chunk of these are due to the way docs ci works itself, that it starts with a specific version of OCaml. In order to do this, the solution must have a specific version of OCaml in it, and this is not always the case, for example, all of the &lt;code&gt;conf-*&lt;/code&gt; packages fail in this way. This particular class of &amp;quot;failures&amp;quot; is not at all important, as mostly they don&apos;t contain useful documentation, but even if they do, if they&apos;re actually being used then they will be built as part of another solution. For example, while &lt;code&gt;conf-faad&lt;/code&gt; fails with this error, the solution of the &lt;code&gt;faad&lt;/code&gt; package pulls it in anyway, so we can build any docs that it includes. Roughly two thirds (685) of the reported failures are due to this, and by checking the resulting HTML docs we can see that we do get docs for 278 of these, so they must be pulled in by other solutions.&lt;/p&gt;
&lt;p&gt;The remaining failures are &amp;quot;real&amp;quot; in the sense that we never currently get docs for these packages. In turn, these can be subcategorised. One class of failures happen with platform-specific packages, for example &lt;code&gt;camlkit&lt;/code&gt; which provides bindings to Cocoa frameworks, and is only available on macOS, or &lt;code&gt;eio_windows&lt;/code&gt; which clearly requires Windows. The current docs-ci setup only builds on Linux, and extending this to other platforms will require a little more work, and is not currently scheduled. These are &amp;quot;fixable&amp;quot; failures.&lt;/p&gt;
&lt;p&gt;The third class of failures are those that will &amp;quot;just never work&amp;quot;. For example, some early versions of &lt;code&gt;domainslib&lt;/code&gt; were released before the OCaml 5.0 APIs were finalised, and so they can only work with alpha versions of OCaml 5.0. We won&apos;t be documenting these.&lt;/p&gt;
&lt;p&gt;Finally there are some more &apos;unexplained&apos; failures, such as &lt;code&gt;docteur.0.0.2&lt;/code&gt;. This one&apos;s particularly interesting as the solve actually succeeds when using the stand-alone tool opam-0install, whereas it&apos;s failing in docs-ci, which uses opam-0install as a library! I&apos;m currently suspicious about the &apos;deprecated&apos; flag, as the failure log contains:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;- git-cohttp-unix -&amp;gt; (problem)
    No usable implementations:
      git-cohttp-unix.3.6.0: Availability condition not satisfied
      git-cohttp-unix.3.5.0: Availability condition not satisfied
      git-cohttp-unix.3.4.0: Availability condition not satisfied
      git-cohttp-unix.3.3.3: Availability condition not satisfied
      git-cohttp-unix.3.3.2: Availability condition not satisfied
      ...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and that flag is the only thing I can immediately see that stands out in &lt;code&gt;git-cohttp-unix&lt;/code&gt;. In contrast, the solution given by opam-0install contains &lt;code&gt;git-cohttp-unix.3.6.0&lt;/code&gt; as a dependency. I suspect fixing this will cause quite a few more packages to succeed.&lt;/p&gt;
&lt;h2&gt;Step 2: building packages&lt;/h2&gt;
&lt;p&gt;The next step, once we&apos;ve got the solutions, is to build the packages. This is using the new method I &lt;a href=&quot;/blog/2025/04/ocaml-docs-ci-and-odoc-3.html&quot;&gt;previously wrote about&lt;/a&gt;. There are about 1,000 packages that fail to build, and once again we can take a look and categorise some of these failures. There are a wider variety of failures here, and it&apos;s quite useful to cross-check with &lt;a href=&quot;https://check.ci.ocaml.org/&quot;&gt;opam health check&lt;/a&gt; to see if it&apos;s known to be broken. Unfortunately OHC only builds the latest versions of everything, so we can&apos;t check in some cases. The interesting issues are where we&apos;re failing to build something that seems to work in OHC.&lt;/p&gt;
&lt;h3&gt;llvm.18&lt;/h3&gt;
&lt;p&gt;This is an interesting type of error, where the build fails because of a missing external dependency. The &lt;code&gt;llvm&lt;/code&gt; package depends upon &lt;code&gt;conf-llvm-static.18&lt;/code&gt;, which should be able to install the depext. Looking at the package, it does indeed have a depext for Debian:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;depexts: [
  [&amp;quot;llvm@18&amp;quot; &amp;quot;zstd&amp;quot;] {os-distribution = &amp;quot;homebrew&amp;quot; &amp;amp; os = &amp;quot;macos&amp;quot;}
  [&amp;quot;llvm-18&amp;quot;] {os-distribution = &amp;quot;macports&amp;quot; &amp;amp; os = &amp;quot;macos&amp;quot;}
  [&amp;quot;llvm-18-dev&amp;quot; &amp;quot;zlib1g-dev&amp;quot; &amp;quot;libzstd-dev&amp;quot;] {os-family = &amp;quot;debian&amp;quot;}
  [&amp;quot;llvm18-dev&amp;quot;] {os-distribution = &amp;quot;alpine&amp;quot;}
  [&amp;quot;llvm18&amp;quot;] {os-family = &amp;quot;arch&amp;quot;}
  [&amp;quot;llvm18-devel&amp;quot;] {os-family = &amp;quot;suse&amp;quot; | os-family = &amp;quot;opensuse&amp;quot;}
  [&amp;quot;llvm18-devel&amp;quot;] {os-distribution = &amp;quot;fedora&amp;quot; &amp;amp; os-version &amp;gt;= &amp;quot;41&amp;quot;}
  [&amp;quot;llvm-devel&amp;quot;] {os-distribution = &amp;quot;fedora&amp;quot; &amp;amp; os-version = &amp;quot;40&amp;quot;}
  [&amp;quot;llvm18-devel&amp;quot; &amp;quot;epel-release&amp;quot;] {os-distribution = &amp;quot;centos&amp;quot;}
  [&amp;quot;devel/llvm18&amp;quot;] {os = &amp;quot;freebsd&amp;quot;}
]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;However, in Debian 12, they&apos;ve already updated to &lt;code&gt;llvm-19&lt;/code&gt;, so the depext is not available.&lt;/p&gt;
&lt;h3&gt;camlimages.5.0.5&lt;/h3&gt;
&lt;p&gt;This one fails due to a linking error. Oddly enough it does seem to work in OHC.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;(cd _build/default &amp;amp;&amp;amp; /home/opam/.opam/4.14/bin/ocamlmklib.opt -g -o freetype/camlimages_freetype_stubs freetype/ftintf.o -ldopt -lfreetype)
# /usr/bin/ld: freetype/ftintf.o: warning: relocation against `Caml_state&apos; in read-only section `.text&apos;
# /usr/bin/ld: freetype/ftintf.o: relocation R_X86_64_PC32 against undefined symbol `Caml_state&apos; can not be used when making a shared object; recompile with -fPIC
# /usr/bin/ld: final link failed: bad value
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;ahrocksdb.0.2.2&lt;/h3&gt;
&lt;p&gt;This one fails in OHC too, but it looks like it&apos;s a build failure with more recent gccs, fixed upstream: https://github.com/ahrefs/ocaml-ahrocksdb/commit/e52316b3d30fededac023141bf8b47da79cabfed&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# run: gcc -O2 -fno-strict-aliasing -fwrapv -fPIC -pthread  -I/usr/include/rocksdb -I /home/opam/.opam/5.3/lib/ocaml -o /tmp/build_02b340_dune/ocaml-configuratordc5e55/c-test-2/test.exe /tmp/build_02b340_dune/ocaml-configuratordc5e55/c-test-2/test.c -lm -lpthread -lrocksdb
# -&amp;gt; process exited with code 1
# -&amp;gt; stdout:
# -&amp;gt; stderr:
#  | In file included from /tmp/build_02b340_dune/ocaml-configuratordc5e55/c-test-2/test.c:4:
#  | /usr/include/rocksdb/version.h:7:10: fatal error: string: No such file or directory
#  |     7 | #include &amp;lt;string&amp;gt;
#  |       |          ^~~~~~~~
#  | compilation terminated.
# Error: discover error
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;alt-ergo.2.2.0&lt;/h3&gt;
&lt;p&gt;Looks like it&apos;s trying to write outside the sandbox. The failure only occurs on alt-ergo 1.3.0 - 2.2.0.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# mkdir -p /home/opam/.opam/4.14/man/man1
# cp -f doc/alt-ergo.1 /home/opam/.opam/4.14/man/man1
# mkdir -p /usr/local/lib/alt-ergo/preludes
# mkdir: cannot create directory &apos;/usr/local/lib/alt-ergo&apos;: Permission denied
# make: *** [Makefile.users:243: install-preludes] Error 1
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;ctypes-foreign.0.18.0&lt;/h3&gt;
&lt;p&gt;This one is a much more interesting failure. The logs show:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;[ERROR] No solution for ctypes-foreign.0.18.0:   * Missing dependency:
            - ctypes-foreign -&amp;gt; ctypes
            unknown package
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;which is happening because of the optimisation I &lt;a href=&quot;/blog/2025/04/ocaml-docs-ci-and-odoc-3.html&quot;&gt;mentioned before&lt;/a&gt; where we build a new &lt;code&gt;opam-repository&lt;/code&gt; with only the packages we&apos;re going to need. In this case, we&apos;ve somehow missed out the &lt;code&gt;ctypes&lt;/code&gt; package. Looking at the opam file for &lt;code&gt;ctypes-foreign&lt;/code&gt;, it has a &lt;code&gt;post&lt;/code&gt; dependency on &lt;code&gt;ctypes&lt;/code&gt;. The &lt;code&gt;post&lt;/code&gt; keyword indicates that &lt;code&gt;ctypes&lt;/code&gt; should be installed with &lt;code&gt;ctypes-foreign&lt;/code&gt;, but that having it as a &amp;quot;normal&amp;quot; dependency would introduce a dependency cycle. Since we require a DAG of dependencies, we explicitly remove any &lt;code&gt;post&lt;/code&gt; dependencies from the set of packages to build, but it seems that &lt;code&gt;opam&lt;/code&gt; would like to know about it anyway!&lt;/p&gt;
&lt;h3&gt;others&lt;/h3&gt;
&lt;p&gt;There are many more. An automatic cross-check with OHC would be really useful, mainly to distinguish between the packages that are broken due to &lt;code&gt;ocaml-docs-ci&lt;/code&gt; issues (like &lt;code&gt;ctypes-foreign&lt;/code&gt;) and those that are broken for other reasons (like &lt;code&gt;ahrocksdb&lt;/code&gt;).&lt;/p&gt;
&lt;h2&gt;Step 3: building docs&lt;/h2&gt;
&lt;p&gt;Finally, we have the actual docs build. This is where we run &lt;code&gt;odoc&lt;/code&gt; and &lt;code&gt;odoc_driver&lt;/code&gt; to produce the HTML docs. All the errors here are ones that we should be able to fix!&lt;/p&gt;
&lt;p&gt;Firstly, there are the internal errors:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Uncaught exception: Failure(&amp;quot;\&amp;quot;rm\&amp;quot; \&amp;quot;-rf\&amp;quot; \&amp;quot;/var/cache/obuilder/merged/582e973685d380d4c91eadc2611eee02c82c5fe4f8bd732e0080fb22bc4404cd\&amp;quot; \&amp;quot;/var/cache/obuilder/work/582e973685d380d4c91eadc2611eee02c82c5fe4f8bd732e0080fb22bc4404cd\&amp;quot; failed with exit status 1&amp;quot;)
2025-05-22 09:30.18: Job failed: Failed: Internal error
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;These are some &lt;code&gt;obuilder&lt;/code&gt; error that needs fixing. Currently we&apos;re just rerunning the job to fix these.&lt;/p&gt;
&lt;h3&gt;odoc.2.0.0&lt;/h3&gt;
&lt;p&gt;Oops, we can&apos;t build our own docs! At least it&apos;s an old version :-)&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;odoc: internal error, uncaught exception:
      File &amp;quot;src/html/link.ml&amp;quot;, line 101, characters 16-22: Assertion failed
      Raised at Odoc_html__Link.href in file &amp;quot;src/html/link.ml&amp;quot;, line 101, characters 16-57
      Called from Odoc_html__Generator.internallink in file &amp;quot;src/html/generator.ml&amp;quot;, line 108, characters 19-49
...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The failure points &lt;a href=&quot;https://github.com/ocaml/odoc/blob/42190737339d9be4510eeeb0e3c47e84badf4d73/src/html/link.ml#L101&quot;&gt;here&lt;/a&gt;, an assertion about the common ancestor of two paths. &lt;a href=&quot;https://github.com/ocaml/odoc/issues/1345&quot;&gt;Issue filed&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;ocaml-base-compiler.4.07.0&lt;/h3&gt;
&lt;p&gt;This one happens because of our &amp;quot;optimisation&amp;quot; to use a base image with OCaml pre-installed. What we &lt;em&gt;actually&lt;/em&gt; do is find the major/minor version of OCaml and use the corresponding docker image - so in this case we&apos;ll use ocaml/opam:debian-12-ocaml-4.07. Now this image actually contains OCaml 4.07.1, and the format of &lt;code&gt;cmt&lt;/code&gt; and &lt;code&gt;cmti&lt;/code&gt; files changed between these releases, so we get a failure.&lt;/p&gt;
&lt;p&gt;We&apos;ll fix this by getting rid of the optimisation and building from an empty switch.&lt;/p&gt;
&lt;h3&gt;lascar.0.7.0&lt;/h3&gt;
&lt;p&gt;This one is quite interesting. It&apos;s another assertion failure in odoc:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;odoc: internal error, uncaught exception:
      File &amp;quot;src/xref2/cpath.ml&amp;quot;, line 364, characters 37-43: Assertion failed
      Raised at Odoc_xref2__Cpath.unresolve_resolved_parent_path in file &amp;quot;src/xref2/cpath.ml&amp;quot;, line 364, characters 37-49
      Called from Odoc_xref2__Cpath.unresolve_module_path in file &amp;quot;src/xref2/cpath.ml&amp;quot;, line 349, characters 28-60
      Called from Odoc_xref2__Tools.fragmap.map_module_decl in file &amp;quot;src/xref2/tools.ml&amp;quot;, line 1792, characters 48-80
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It&apos;s happening when we &apos;unresolve&apos; a previously resolved path. We end up having to do this when something about the path has changed, in this case while we&apos;re handling a &lt;code&gt;S with module Foo = Bar&lt;/code&gt; or similar. Issue &lt;a href=&quot;https://github.com/ocaml/odoc/issues/1346&quot;&gt;filed&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;camlp5&lt;/h3&gt;
&lt;p&gt;This one actually occurs in &lt;code&gt;odoc_driver&lt;/code&gt; rather than in &lt;code&gt;odoc&lt;/code&gt; itself.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;odoc_driver_voodoo: [DEBUG] Found cmi_only_lib in dir: /home/opam/.opam/4.08/lib/camlp5
odoc_driver_voodoo: internal error, uncaught exception:
                    Invalid_argument(&amp;quot;\&amp;quot;/home/opam/.opam/4.08/lib/camlp5\&amp;quot;: invalid segment&amp;quot;)
                    
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here we&apos;re trying to add a segment to a path, but rather than a single path segment we&apos;ve got an entire fully qualified path. Issue &lt;a href=&quot;https://github.com/ocaml/odoc/issues/1347&quot;&gt;filed&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;It&apos;s pretty good that we&apos;ve only got 4 types of error happening at the doc-generation phase. However, as a whole, any error that occurs earlier in the pipeline ends up with a missing documentation tab on the website, and we need to do a bit more so that the actual problem can be tracked down and fixed. This is obviously a more general problem than just the docs, and one that &lt;a href=&quot;https://check.ci.ocaml.org&quot;&gt;opam health check&lt;/a&gt; seeks to highlight. However, the current incarnation of OHC is significantly less efficient than docs-ci, so generalising the approach we&apos;ve taken with &lt;a href=&quot;https://github.com/jonludlam/opamh&quot;&gt;opamh&lt;/a&gt; should really help with making this more responsive.&lt;/p&gt;
&lt;p&gt;In addition, a number of the issues seen here could be addressed with a tool my colleague &lt;a href=&quot;https://ryan.freumh.org/&quot;&gt;Ryan&lt;/a&gt; is working on: &lt;a href=&quot;https://ryan.freumh.org/enki.html&quot;&gt;Enki&lt;/a&gt;. This tool would allow us to run a solve that actually determines not only the set of packages we wish to install, but the platform to install onto - e.g. for &lt;code&gt;eio_windows&lt;/code&gt; the solution would be to install on Windows, and for &lt;code&gt;llvm.18-static&lt;/code&gt; the solution might be Fedora 40.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/05/docs-progress.html" rel="alternate" title="Progress in OCaml docs"/>
    <category term="odoc"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/05/lots-of-things.html</id>
    <title type="text">Lots of things have been happening</title>
    <updated>2025-05-20T00:00:00Z</updated>
    <published>2025-05-20T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;I&apos;ve been working on a whole lot of thing recently in many different areas, making what&apos;s felt like only a bit of progress in each. Consequently I&apos;ve not felt like I had anything substantial to say, so I haven&apos;t written up anything for a while.&lt;/p&gt;
&lt;p&gt;Time for a little summary of things then!&lt;/p&gt;
&lt;h2&gt;Ocaml-docs-ci&lt;/h2&gt;
&lt;p&gt;I&apos;ve been working with &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark Elvers&lt;/a&gt; on getting the docs CI running using Odoc 3.0. There are quite a few changes involved, both in how we&apos;re &lt;a href=&quot;/blog/2025/04/ocaml-docs-ci-and-odoc-3.html&quot;&gt;building the packages&lt;/a&gt; but also how we&apos;re running odoc - it&apos;s building using &lt;code&gt;odoc_driver&lt;/code&gt; rather than &lt;code&gt;voodoo&lt;/code&gt; now, and while it&apos;s looking promising now we had hit a few hurdles along the way.&lt;/p&gt;
&lt;p&gt;We set the CI going last weekend but discovered that it was having some issues building packages using OCaml 5.3.0. The way the builds work is that we first do a &amp;quot;solve&amp;quot; step for each version of every package so we&apos;ve got exact versions of all of the packages required to build them. We then look through that solution to figure out the version of OCaml required, and the (to avoid a little bit of work) we start from one of the &lt;a href=&quot;https://hub.docker.com/r/ocaml/opam&quot;&gt;opam docker images&lt;/a&gt; for that version of OCaml.&lt;/p&gt;
&lt;p&gt;When installing a package using opam it does a few operations that scale with the size of the opam repository, which ends up adding around ten of seconds to the build time. When we&apos;re building 50,000 packages, this adds up to quite a lot of time, so we short-cut this process with the simple expedient of creating an opam-repository that only contains the packages we need for the build. However, since we&apos;ve already got a few packages installed in the image, we need to make sure our repository contains these packages too, otherwise opam gets thoroughly confused. My mistake was that we were missing out the `ocaml-compiler` package, which is new in OCaml 5.3.0, which led to the builds failing. Adding this in and kicking off the build again it&apos;s now got a lot further - at time of writing it has built 14,000 packages, there are 6,000 still building, and 1000 that have failed. If it continues in a similar fashion, this will compare quite favourably with the docs CI that&apos;s currently powering ocaml.org, where it has successfully built 17,000 packages, and 4,500 have failed.&lt;/p&gt;
&lt;p&gt;Mark has been working on a different approach to the build process, which is to come up with a new binary that doesn&apos;t do any of the &lt;code&gt;O(n)&lt;/code&gt; operations and just builds the package! This is definitely a promising direction, and I&apos;m hoping to take a look at &lt;a href=&quot;https://github.com/mtelvers/ohc&quot;&gt;his prototype&lt;/a&gt; soon.&lt;/p&gt;
&lt;p&gt;Meanwhile, &lt;a href=&quot;https://choum.net&quot;&gt;panglesd&lt;/a&gt; is working on integrating this into the ocaml.org site, and is making good progress. He spotted last week that we were overwriting the `status.json` file that comes out of `odoc_driver` which we will use to power the redirections on ocaml.org. One of the changes of odoc 3.0 is that we carefully put modules into a directory structure that represents the library in which they are found. It&apos;s long been a pain that OCaml libraries (what Ocamlfind unhelpfully calls &apos;packages&apos;) are not always the same name as the opam package in which they&apos;re found. For example, the package &lt;code&gt;ocamlfind&lt;/code&gt; contains the library &lt;code&gt;findlib&lt;/code&gt;. So to help the user figure out where to find the module, we&apos;re putting it into the URL of the docs, and therefore into the breadcrumbs. The downside is that the modules are now in a different place on the website to where they were before, so the &lt;code&gt;status.json&lt;/code&gt; file is there to help with the redirections.&lt;/p&gt;
&lt;h2&gt;Notebooks&lt;/h2&gt;
&lt;p&gt;I&apos;ve been working on Merlin integration with the notebooks, which has been a fun little project. The bits that needed improving most were that merlin didn&apos;t work with toplevel-style code, and that each cell was a separate typing context, so while you could define a function in one cell and execute it in another, Merlin would tell you the function was undefined.&lt;/p&gt;
&lt;p&gt;For the toplevel-style code, what I&apos;ve ended up doing is to essentially strip out all of the toplevel bits and pieces, and replace them with whitespace. So where I have a cell that looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# let x = 1 + 2;;
  - val x : int = 3
# let y = x + 1;;
  - val y : int = 4
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I tell Merlin that the contents are:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;  let x = 1 + 2;;

  let y = x + 1;;
               
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where I&apos;m careful to maintain the position of the original code. This bit is working quite nicely, but only when the code is syntactically correct, as I&apos;m using the standard toplevel parser to figure out where the expression ends. I think I&apos;m going to end up needing to write a custom parser for this, so something that will end on a &lt;code&gt;;;&lt;/code&gt; but ignore them in string constants, comments and so on.&lt;/p&gt;
&lt;p&gt;The approach I&apos;ve taken for the second problem is to treat each cell as a separate module. I then write out a &lt;code&gt;cmi&lt;/code&gt; file into the virtual filesystem as &lt;code&gt;cell__id_0.cmi&lt;/code&gt; and &lt;code&gt;open&lt;/code&gt; all the previous modules in &apos;line 0&apos; of every cell. I then remap all of the reported locations by removing &apos;line 0&apos;.&lt;/p&gt;
&lt;p&gt;There are a number of issues with the current approaches: 1. The stripping of the toplevel bits is a little fragile, and currently only works when the toplevel is syntactically correct. This is fairly fixable. 2. When the contents of the cells change we need to flush any caches merlin and the compiler have. 3. An &lt;code&gt;open&lt;/code&gt; statement in once cell does _not_ cause the module to be available in the next cell. 4. A lot of cells leads to a lot of opens!&lt;/p&gt;
&lt;p&gt;I suspect that this the &apos;cells as modules&apos; approach might end up being a bit of a dead-end, so I&apos;ll have a chat with &lt;a href=&quot;https://github.com/voodoos&quot;&gt;Ulysse&lt;/a&gt; to figure out the next experiment.&lt;/p&gt;
&lt;h2&gt;Oxcaml&lt;/h2&gt;
&lt;p&gt;I&apos;ve been working on trying out oxcaml too, which has been a bit challenging. Firstly, although Jane Street provide a version of &lt;code&gt;js_of_ocaml&lt;/code&gt;, the toplevel didn&apos;t work. Fortunately, my amazing colleagues &lt;a href=&quot;https://patrick.sirref.org/&quot;&gt;Patrick O&apos;Ferris&lt;/a&gt; and &lt;a href=&quot;https://github.com/art-w&quot;&gt;Arthur Wendling&lt;/a&gt; spent a good chunk of time fixing this and provided an &lt;a href=&quot;https://github.com/patricoferris/opam-repository-js#with-extensions&quot;&gt;opam repository&lt;/a&gt; with the relevant changes, without which I would have not been able to make any progress. Thanks, guys! So my goal of making my notebooks work with it looked doable, but I almost immediately hit more dependency issues that make it problematic to port the whole site over - including odoc and various PPXes that I use.&lt;/p&gt;
&lt;p&gt;I&apos;ve therefore decided that I would bring forward a feature that I&apos;d had in mind for a while - that we could have different &amp;quot;backends&amp;quot; for the notebooks. So I&apos;d still build the frontend using &amp;quot;normal&amp;quot; OCaml, but the web-worker serving as the toplevel would be an oxcaml one.&lt;/p&gt;
&lt;p&gt;Of course, it didn&apos;t work first time! After a bit of head-scratching, it turned out that the interface between the worker and the main thread, although I&apos;d &lt;em&gt;almost&lt;/em&gt; got it ocaml-agnostic, wasn&apos;t quite right. The way it works is that it uses the jsonrpc protocol to communicate, and while it had marshalled the requests into a string, it hadn&apos;t turned that final OCaml string into a Javascript string, so it was sending the js_of_ocaml representation of a string as an object, rather than a simple string. When the frontend and workers were built with different compilers, this ended up just failing with an obscure error, which took a good deal of time to track down. Once that was fixed, it was just a case of making sure I could have 2 independent &apos;switches&apos; on my site - one for oxcaml and one for standard OCaml.&lt;/p&gt;
&lt;p&gt;The upshot of all this is that I now have a semi-working version of the notebooks using oxcaml. As an initial demonstration I ported one of my colleague &lt;a href=&quot;https://github.com/cuihtlauac&quot;&gt;Cuihtlauac&lt;/a&gt;&apos;s oxcaml tutorial docs to the notebook format, and it &lt;a href=&quot;/notebooks/oxcaml/local.html&quot;&gt;works quite nicely&lt;/a&gt;.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/05/lots-of-things.html" rel="alternate" title="Lots of things have been happening"/>
    <category term="odoc"/>
    <category term="ocaml"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/05/ticks-solved-by-ai.html</id>
    <title type="text">Solving First-year OCaml exercises with AI</title>
    <updated>2025-05-07T00:00:00Z</updated>
    <published>2025-05-07T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;My colleague &lt;a href=&quot;https://toao.com&quot;&gt;Sadiq Jaffer&lt;/a&gt; and I have been working on a little project to see how well small AI models can solve the OCaml exercises we give to our first-year students at the University of Cambridge. Sadiq has done an excellent &lt;a href=&quot;https://toao.com/blog/ocaml-local-code-models&quot;&gt;write up&lt;/a&gt; of our initial results, which you should all go and read! The tl;dr though, as Sadiq writes, is that even some of the smaller models would score top marks on these exercises!&lt;/p&gt;
&lt;p&gt;One interesting aspect we discovered quite quickly is that we had to make the testing feedback a little more generous than just &amp;quot;exception raised&amp;quot;! The problems are presented as a Jupyter notebook using &lt;a href=&quot;https://github.com/akabe&quot;&gt;akabe&apos;s&lt;/a&gt; excellent OCaml kernel, with &lt;a href=&quot;https://nbgrader.readthedocs.io/en/stable/&quot;&gt;nbgrader&lt;/a&gt; to do the assessment. Our students can see the tests that are run, and if they fail they&apos;re able to copy the test cell out and play with their code to figure out exactly what went wrong. The AI models, however, have a far less interactive experience, and get just 3 chances to write code that passes the tests. We found that the performance of the models increased hugely when we adjusted the test cells such that they clearly indicated which test failed, the results that were expected, and the results the code actually produced.&lt;/p&gt;
&lt;p&gt;Of course, we &lt;a href=&quot;https://anil.recoil.org/notes/claude-copilot-sandbox&quot;&gt;already knew&lt;/a&gt; that AI models can code OCaml very well, and we (along with the rest of the teaching world) are still ruminating on the implications of this from a pedagogical perspective. Our plan, though, is to try and make the &apos;problem&apos; worse by training these models on more OCaml code, and see just how well we can get them to perform! It&apos;s pretty amazing, and a little startling to know that a model that&apos;ll run pretty comfortably on my laptop can solve these problems so well even without extra training, though given how hot it gets, I&apos;d rather not have the laptop on my actual lap while it&apos;s doing so!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/05/ticks-solved-by-ai.html" rel="alternate" title="Solving First-year OCaml exercises with AI"/>
    <category term="ai"/>
    <category term="teaching"/>
    <category term="ocaml"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/05/oxcaml-gets-closer.html</id>
    <title type="text">OxCaml is getting closer...</title>
    <updated>2025-05-02T00:00:00Z</updated>
    <published>2025-05-02T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;I joined the OxCaml weekly meeting representing Tarides for the first time this week, as Jane Street gear up to an official release of their OxCaml compiler.&lt;/p&gt;
&lt;p&gt;It seems that mainly what needs to be done before the release can be made is to ensure there is some reasonable documentation for the new features, and that a reasonable number of packages are working, so people are furiously writing and bugfixing to try and get this ready.&lt;/p&gt;
&lt;p&gt;As well as this though, there are some challenges of a more organisational level that will need to be addressed to ensure the success of the project. Jane Street have long had a public branch of their compiler, but while they&apos;ve had patches internally to ensure the tooling and other libraries work, these patches haven&apos;t previously been made public in a usable way. In order for OxCaml to be useful, it will clearly need these patches not only to be available, but also to be maintained and to easily allow contributions from the community -- in short, they need to be properly Open Source!&lt;/p&gt;
&lt;p&gt;Personally, I&apos;m looking forward to seeing their branch of &lt;a href=&quot;https://ocaml.github.io/odoc/&quot;&gt;odoc&lt;/a&gt; and having a look to see how the modes will fit into the documentation. I&apos;m also keen to see whether the &lt;a href=&quot;/blog/2025/04/this-site.html&quot;&gt;notebook features&lt;/a&gt; I&apos;ve been working on can be ported over to run on OxCaml!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/05/oxcaml-gets-closer.html" rel="alternate" title="OxCaml is getting closer..."/>
    <category term="ocaml"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/05/ai-for-climate-and-nature-day.html</id>
    <title type="text">AI for Climate &amp; Nature Community Day</title>
    <updated>2025-05-01T00:00:00Z</updated>
    <published>2025-05-01T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;&lt;img src=&quot;./melissa.jpg&quot; alt=&quot;Melissa Leach&quot; &gt;
&lt;em&gt;Melissa Leach introducing the day&lt;/em&gt; Today was the &amp;quot;AI for Climate &amp;amp; Nature Community Day&amp;quot; at the &lt;a href=&quot;https://map.cam.ac.uk/?maplon=0.12032&amp;amp;maplat=52.20354&amp;amp;mapzoom=18&amp;amp;maplayers=Building+Labels%2CExternal+Sites%2CColleges%2CUniversity+Sites%2CBuildings%2CTransport&amp;amp;mapfeature=mfid257%2CBuildings&quot;&gt;David Attenborough Building&lt;/a&gt;. A whole bunch of the EEG were either presenting or contributing in some way so I thought I&apos;d come along to see what&apos;s going on.&lt;/p&gt;
&lt;h2&gt;Keynote and main talks&lt;/h2&gt;
&lt;p&gt;Following the intro talks from Professors &lt;a href=&quot;https://www.cambridgeconservation.org/about/people/prof-melissa-leach/&quot;&gt;Melissa Leach&lt;/a&gt; and &lt;a href=&quot;https://www.zoo.cam.ac.uk/directory/bill-sutherland&quot;&gt;Bill Sutherland&lt;/a&gt;, the day started with the keynote talk from &lt;a href=&quot;https://www.biology.ox.ac.uk/people/amy-hinsley&quot;&gt;Amy Hinsley&lt;/a&gt;, who, using the specific case of animial trafficking, talked about the need to make AI in conservation equitable, explainable and useful.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./amy.jpg&quot; alt=&quot;Amy Hinsley&quot; &gt;
&lt;em&gt;Amy Hinsley delivering the keynote talk&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We then moved into the first session with &lt;a href=&quot;https://www.geog.cam.ac.uk/people/lines/&quot;&gt;Emily Lines&lt;/a&gt; from the Geography Department who talked about the challenges processing sensor data in the context of forests. Her group has a variety of data collected from forests across Europe, collected from using many different methods, from drones taking pictures of the canopies to ground-based laser scanners producing 3d point clouds. The challenge is then not only to identify individual trees, which is pretty tricky, but also to then distinguish between the leaves of the trees and the wood.&lt;/p&gt;
&lt;p&gt;After Emily came &lt;a href=&quot;https://ai.cam.ac.uk/people/robert-rouse.html&quot;&gt;Robert Rouse&lt;/a&gt; from the &lt;a href=&quot;https://iccs.cam.ac.uk&quot;&gt;ICCS&lt;/a&gt;, who&apos;s using a small neural net and genetic algorithms to extend a study from the RSPB on figuring out an optimal way to do some land use adjustments to cut carbon and improve outcomes for birds, whilst not significantly impacting the ability to produce food.&lt;/p&gt;
&lt;p&gt;We then had &lt;a href=&quot;https://www.zoo.cam.ac.uk/directory/dr-sam-reynolds&quot;&gt;Sam Reynolds&lt;/a&gt; and &lt;a href=&quot;https://toao.com&quot;&gt;Sadiq Jaffer&lt;/a&gt; who talked about their project; using AI to sift through millions of papers looking for those relevant to a specified conservation topic. They&apos;re able to directly compare their results with results obtained by manually doing this process, a project that&apos;s been going on over the last 20 or so years summing to something like 75 man years of effort. In the end they only missed a few papers that the manual process had found, but actually found many relevant papers that had been missed - and all in only a few days of compute.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;./sadiq.jpg&quot; alt=&quot;Sam Reynolds and Sadiq Jaffer&quot; &gt;
&lt;em&gt;Sam Reynolds and Sadiq Jaffer sorting millions of papers&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;Lightning talks&lt;/h2&gt;
&lt;p&gt;We then had a number of &apos;lightning talks&apos;, with each presenter having only three minutes to talk about their work.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.maths.cam.ac.uk/person/ss3299&quot;&gt;Sebastian Schemm&lt;/a&gt; presented his work on creating a foundational model for the climate&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.eng.cam.ac.uk/profiles/ac685&quot;&gt;Alice Cicirello&lt;/a&gt; talked about the prospects of applying machine learning to &lt;a href=&quot;https://en.wikipedia.org/wiki/Marine_cloud_brightening&quot;&gt;Marine Cloud Brightening&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.maths.cam.ac.uk/person/sdat2&quot;&gt;Simon Thomas&lt;/a&gt; has been looking at analysing the heights of tropical cyclone storm surges&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/niccolozanotti&quot;&gt;Niccolò Zanotti&lt;/a&gt; gave us an introduction to &lt;a href=&quot;https://github.com/cambridge-ICCS/FTorch&quot;&gt;FTorch&lt;/a&gt;, a library to integrate the worlds of PyTorch and Fortran&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.nceo.ac.uk/contact-us/people/dr-simon-driscoll/&quot;&gt;Simon Driscoll&lt;/a&gt; then talked about melt ponds on arctic sea ice, a poorly understood but important component of the climate in the Arctic.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.zoo.cam.ac.uk/directory/emilio-luz-ricca&quot;&gt;Emilio Luz-Ricca&lt;/a&gt; talked about his project to apply machine learning to predict hunting pressure&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://orlando-code.github.io&quot;&gt;Orlando Timmerman&lt;/a&gt; gave us some insights into how he&apos;s been using machine learning to predict the future of coral reefs, and how we might use this to help with their conservation.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.zoo.cam.ac.uk/directory/ruari-marshall-hawkes&quot;&gt;Ruari Marshall-Hawkes&lt;/a&gt; showed us how to listen very carefully to figure out population numbers,&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.linkedin.com/in/harriet-branson-a93a8313b/&quot;&gt;Hattie Branson&lt;/a&gt; from &lt;a href=&quot;https://www.fauna-flora.org&quot;&gt;Fauna &amp;amp; Flora&lt;/a&gt; talked about habitat detection in South Sudan,&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.linkedin.com/in/martakoch/&quot;&gt;Marta Koch&lt;/a&gt; showed us an analysis of how well ChatGPT, Claude and the like would perform at setting the agendas for SDPs,&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.linkedin.com/in/zhengpeng-feng-2410a132a/&quot;&gt;Frank Feng&lt;/a&gt; finished the session with a talk on the &lt;a href=&quot;https://www.cst.cam.ac.uk/seminars/list/227335&quot;&gt;Barlow Twins Earth Foundation Model&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;div style=&quot;display: grid; grid-template-columns: 1fr 1fr; gap: 20px;&quot;&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;sebastian.jpg&quot; alt=&quot;Sebastian Schemm&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Sebastian Schemm&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;alice.jpg&quot; alt=&quot;Alice Cicirello&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Alice Cicirello&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;simon.jpg&quot; alt=&quot;Simon Thomas&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Simon Thomas&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;simond.jpg&quot; alt=&quot;Simon Driscoll&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Simon Driscoll&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;emilio.jpg&quot; alt=&quot;Emilio Luz-Ricca&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Emilio Luz-Ricca&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;orlando.jpg&quot; alt=&quot;Orlando Timmerman&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Orlando Timmerman&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;ruari.jpg&quot; alt=&quot;Ruari Marshall-Hawkes&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Ruari Marshall-Hawkes&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;hattie.jpg&quot; alt=&quot;Hattie Branson&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Hattie Branson&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure style=&quot;margin:0; width: 100%;&quot;&gt;
        &lt;img src=&quot;frank.jpg&quot; alt=&quot;Frank Feng&quot; style=&quot;max-width: 100%; height: auto;&quot;&gt;
        &lt;figcaption&gt;Frank Feng&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/div&gt;
&lt;h2&gt;Discussions&lt;/h2&gt;
&lt;p&gt;We then split up into three discussion groups; one on the future of this work, one on how to continue building this community of researchers, and the last on applying AI to real-world problems. As a newcomer to the field I was interested in the direction it&apos;s heading in, so I joined in &lt;a href=&quot;https://dorchard.github.io&quot;&gt;Dominic Orchard&lt;/a&gt;&apos;s led session on the future of AI.&lt;/p&gt;
&lt;p&gt;We had a fascinating discussion on both the immediate things we can do and longer term worries. We were imagining a world where AI becomes &apos;just a tool&apos; that we don&apos;t need to be experts in to apply it, but right now we&apos;re in a much more tightly coupled collaborative world where we need experts in AI to complement the experts in the application field to make progress. This comes with challenges - applying for funding for multidisciplinary work is not the norm, so we spent some time discussing this too.&lt;/p&gt;
&lt;p&gt;One of our group spoke about statistics now being &apos;just a tool&apos;, but it&apos;s been one that we&apos;ve worked with for a long time now and we know where the sharp corners are. We have protocols for applying statistical tools and we have diagnostic plots to tell us whether the results are trustworthy, but not only do we not have these for AI models, but it&apos;s not yet clear whether such a thing will be even possible.&lt;/p&gt;
&lt;p&gt;Overall it was a fascinating day, and I&apos;m very much looking forward to following the work of these outstanding researchers, and maybe even contributing to their work in some way in the future.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Thanks to &lt;a href=&quot;https://anil.recoil.org&quot;&gt;Anil Madhavapeddy&lt;/a&gt; for the photos of the day.&lt;/em&gt;&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/05/ai-for-climate-and-nature-day.html" rel="alternate" title="AI for Climate &amp; Nature Community Day"/>
    <category term="ai"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/04/ocaml-docs-ci-and-odoc-3.html</id>
    <title type="text">OCaml-Docs-CI and Odoc 3</title>
    <updated>2025-04-29T00:00:00Z</updated>
    <published>2025-04-29T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;The release of Odoc 3 means that we need to update the &lt;a href=&quot;https://docs.ci.ocaml.org&quot;&gt;docs-ci&lt;/a&gt; project so that the documentation that appears on &lt;a href=&quot;https://ocaml.org/p/&quot;&gt;ocaml.org&lt;/a&gt; is using the latest, greatest Odoc. With this major release of Odoc, it&apos;s also time to give the CI pipeline a bit of an overhaul too, and fix some of the irritations that it causes.&lt;/p&gt;
&lt;h2&gt;The challenge of documenting OCaml&lt;/h2&gt;
&lt;p&gt;As I wrote about &lt;a href=&quot;/blog/2025/04/semantic-versioning-is-hard.html&quot;&gt;recently&lt;/a&gt;, the APIs of OCaml libraries are dependent not only on the version of its package, but possibly also on the versions of any of its dependencies. Due to this fact, to produce the docs for ocaml.org means that sometimes we need to build the docs for a particular version of a particular package multiple times with different versions of its dependencies.&lt;/p&gt;
&lt;p&gt;It&apos;s clearly impractical to try to build every possible combination, so what we do is to run the opam solver once for each version of each package. This gives us a set of packages at particular versions. We then take that, and for each package in the set, we pluck out &lt;em&gt;its&lt;/em&gt; dependencies from the set, producing a &amp;quot;universe&amp;quot; of dependencies for every package in the set. Let&apos;s look at a very simple example; the package &lt;code&gt;cry&lt;/code&gt; from the &lt;a href=&quot;https://www.liquidsoap.info&quot;&gt;LiquidSoap&lt;/a&gt; project.&lt;/p&gt;
&lt;p&gt;The oldest version of &lt;code&gt;cry&lt;/code&gt; from before the &lt;a href=&quot;https://discuss.ocaml.org/t/opam-repository-archival-phase-1-unavailable-packages/15797/6&quot;&gt;Great Purge&lt;/a&gt; was 0.2.2, which when solved produced the following dependencies:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;cry.0.2.2
ocaml.4.05.0
ocaml-base-compiler.4.05.0
ocaml-config.1
ocamlfind.1.9.6
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and the oldest version of &lt;code&gt;cry&lt;/code&gt; after the purge is 0.6.0 which produces the following solution:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;cry.0.6.0
ocaml.5.2.1
ocaml-base-compiler.5.2.1
ocaml-config.3
ocamlfind.1.9.6
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;so we we can see from these two solutions that we&apos;ll need to build &lt;code&gt;ocamlfind.1.9.6&lt;/code&gt; twice, once with &lt;code&gt;ocaml.4.05.0&lt;/code&gt; and once with &lt;code&gt;ocaml.5.2.1&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Once we&apos;ve got, for every version of every package, a set of dependency universes, we choose one of these to be the one presented to the user under the &lt;code&gt;ocaml.org/p/&lt;/code&gt; hierarchy. For all of the other universes, we build the package againt them, and put the docs under the &lt;code&gt;ocaml.org/u/&lt;/code&gt; hierarchy.&lt;/p&gt;
&lt;h2&gt;Performing the builds&lt;/h2&gt;
&lt;p&gt;Once we&apos;ve got a complete set of solutions and builds to do, the current CI pipeline batches the builds up to try and build as many packages as possible in as few builds as possible. While this works well enough, it does mean that we build a lot packages more than once - dune, for example, is built thousands of times during this process, producing exactly the same binaries each time.&lt;/p&gt;
&lt;p&gt;In the new pipeline, I wrote a &lt;a href=&quot;https://github.com/jonludlam/opamh&quot;&gt;little tool&lt;/a&gt; that allows opam packages to be archived and restored, which happens to work nicely because we&apos;re always building the packages in the same container in the same location. This means there are no worries about relocatability, although that is something that is &lt;a href=&quot;https://www.dra27.uk/blog/platform/2025/04/22/branching-out.html&quot;&gt;nearly here!&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The downside to this is that our storage requirements are quite a bit larger, as we&apos;re storing the entire package rather than just the bits that odoc needs. However, we were always going to use more storage than before simply because the new &lt;code&gt;odoc&lt;/code&gt; and &lt;code&gt;odoc_driver&lt;/code&gt; pair are more capable, and the new features like &lt;a href=&quot;https://github.com/ocaml/odoc/pull/909&quot;&gt;source code rendering&lt;/a&gt; and &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1121/files#diff-10c8829023814c0bcc3316f95f643623404c000b13c68849ef3d61097a6e03a6R1-R415&quot;&gt;classify&lt;/a&gt; require more files from the original package.&lt;/p&gt;
&lt;p&gt;The upshot is that I&apos;ll be working with &lt;a href=&quot;https://www.tunbury.org/&quot;&gt;Mark Elvers&lt;/a&gt; to move the docs CI from its current machine to a shiny new &lt;a href=&quot;https://www.tunbury.org/blade-reallocation/&quot;&gt;blade server&lt;/a&gt;.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/04/ocaml-docs-ci-and-odoc-3.html" rel="alternate" title="OCaml-Docs-CI and Odoc 3"/>
    <category term="odoc"/>
    <category term="docs-ci"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/04/odoc-3.html</id>
    <title type="text">Odoc 3: So what?</title>
    <updated>2025-04-25T00:00:00Z</updated>
    <published>2025-04-25T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Odoc 3 was &lt;a href=&quot;https://discuss.ocaml.org/t/ann-odoc-3-0-released/16339&quot;&gt;released last month&lt;/a&gt; and although we did write a list of the new features, I don&apos;t think we&apos;ve made it clear enough why anyone should care.&lt;/p&gt;
&lt;p&gt;It&apos;s &lt;strong&gt;manuals&lt;/strong&gt;, the theme of Odoc 3 is &lt;strong&gt;manuals&lt;/strong&gt;. It&apos;s got a load of features to make it much better for writing &lt;code&gt;mld&lt;/code&gt; pages (files written using odoc&apos;s markup) to document your packages and their relationship to the surrounding ecosystem. Previous versions of Odoc were very library-centric, in that while we did have mld-file support, most of the effort went into making sure that we were generating correct per-module pages, which show the shape of your API even if you&apos;ve not put in any doc comments at all. We&apos;ve still got that, obviously, but we&apos;ve added many features to make write &lt;code&gt;mld&lt;/code&gt; pages far more useful, and we&apos;re really hoping that these will draw people in to make documenting packages a much more enjoyable experience.&lt;/p&gt;
&lt;h2&gt;Odoc&apos;s special skill: links!&lt;/h2&gt;
&lt;p&gt;But why you might want to use Odoc at all for your package&apos;s manuals, rather than, say, markdown, asciidoc, rst or any other similar language? The biggest thing that Odoc brings, and has always brought, is &lt;strong&gt;reliable linking&lt;/strong&gt;. Just write &lt;code&gt;{!Module.func}&lt;/code&gt; and Odoc will check that the target exists and ensure that the link goes to the correct place, no matter how complex the definition of &lt;code&gt;Module&lt;/code&gt; is or what the layout of the docs. We can link to almost all elements of an OCaml library, from modules and types through to fields of records, exceptions and extensions, and we have facilities for disambiguating, so if you happen to have both a module &lt;code&gt;S&lt;/code&gt; and a module type &lt;code&gt;S&lt;/code&gt; you can easily link to whichever you please.&lt;/p&gt;
&lt;p&gt;In Odoc 2 though, these links were pretty limited - the only ones possible were only those to docs and API elements (modules, types, values, etc) in your own package, or to API elements in any libraries that your package depends on. When writing API docs, which tend to be at the level of types and functions, this wasn&apos;t a huge problem, but when considering manuals this turned out to be a really limiting constraint. For example, in Odoc&apos;s own docs, we really want to have a link to &lt;code&gt;odoc-driver&lt;/code&gt;, but since &lt;code&gt;odoc-driver&lt;/code&gt; is a separate package and depends upon &lt;code&gt;odoc&lt;/code&gt;, the only way to do that in Odoc 2.x would be to use an HTML link. With Odoc 3, this constraint is gone, so you can &lt;strong&gt;link to any other package or library&lt;/strong&gt;. The link to &lt;code&gt;odoc-driver&lt;/code&gt; would look like &lt;code&gt;{!/odoc-driver/page-index}&lt;/code&gt;, as can be seen in &lt;a href=&quot;https://github.com/ocaml/odoc/blob/master/doc/driver.mld#L10&quot;&gt;odoc&apos;s source&lt;/a&gt;. The only requirement is that you must be able to simultaneously install all of the packages you&apos;d like to link to, so you can&apos;t easily link to, for example, different versions of the same package.&lt;/p&gt;
&lt;p&gt;This will be particularly useful for any projects that&apos;s grouped into multiple packages. For example, the &lt;a href=&quot;https://mirage.io&quot;&gt;Mirage project&lt;/a&gt;. The main package there -- &lt;code&gt;mirage&lt;/code&gt; -- is actually right at the bottom of the dependency hierarchy, but it&apos;s the perfect place to have docs that link to all of the other Mirage packages. On a smaller scale, the &lt;a href=&quot;https://github.com/ocaml-multicore/picos&quot;&gt;Picos project&lt;/a&gt; consists of multiple packages all from a single git repository, and this would allow the docs pages from the &lt;code&gt;picos&lt;/code&gt; package to link to any of the other packages.&lt;/p&gt;
&lt;p&gt;Of course there are also a lot of other new features in this release, which are called out in the &lt;a href=&quot;https://discuss.ocaml.org/t/ann-odoc-3-beta-release/16043&quot;&gt;annoucement post on discuss&lt;/a&gt;, some of which I may post about in the future.&lt;/p&gt;
&lt;h2&gt;Can I use it now?&lt;/h2&gt;
&lt;p&gt;Of course! These new features can be used right now, so long as you&apos;re happy to self-host the docs. All that&apos;s needed is to create a switch containing all the packages you&apos;re interested in together, and use &lt;code&gt;odoc_driver&lt;/code&gt; to generate the HTML and push them to your web server. At time of writing though, ocaml.org is still using Odoc 2.4, so any packages that are published to opam that choose to use these new features will be missing the new features. Furthermore, it&apos;s actually quite a challenge to do this, since we&apos;ll have to extend the package-universe solutions to include all relevant packages, for which we need extra fields in the opam metadata.&lt;/p&gt;
&lt;h2&gt;What&apos;s next?&lt;/h2&gt;
&lt;p&gt;We&apos;re actively working on getting Odoc 3 into the pipeline generating the docs found in https://ocaml.org/p/. This will bring with it some of the developments that landed in Odoc 2, but didn&apos;t make it onto ocaml.org - for example, the rendering of source pages. Not only are there challenges related to the package-universe solutions as mentioned above, but the storage requirements are considerably larger, so I&apos;ll be working with &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark Elvers&lt;/a&gt; to complete this project.&lt;/p&gt;
&lt;p&gt;We&apos;ve also got work to do to update the build rules in dune to take advantage of these features. While &lt;code&gt;odoc_driver&lt;/code&gt; works very well as part of the process of deploying a docs site, it&apos;s quite impractical as a tool to help while you&apos;re actually writing the docs. For that, we&apos;ll need to make sure &lt;code&gt;dune&lt;/code&gt; understands how to use these new features. Fortunately we&apos;ve had some experience with those rules in the past, and part of the work that&apos;s gone into Odoc 3 was to ensure that incremental build rules should be far more straightforward to write than for Odoc 2. In addition, some of the logic that previously only existed in &lt;a href=&quot;https://github.com/ocaml-doc/voodoo&quot;&gt;Voodoo&lt;/a&gt; - the old driver that builds docs for ocaml.org - has been integrated into &lt;code&gt;odoc&lt;/code&gt; itself, meaning one again that getting dune to produce correct docs for non-dune packages (e.g. the standard library!) should again be simpler.&lt;/p&gt;
&lt;p&gt;After we&apos;ve done these, there are plans afoot to make more improvements to the manual writing experience. &lt;a href=&quot;https://choum.net/&quot;&gt;@panglesd&lt;/a&gt; has been investigating how to add admonitions to the language, I&apos;ve been thinking about custom tag support, we&apos;re looking at the &lt;a href=&quot;https://discuss.ocaml.org/t/ann-oxidizing-ocaml-an-update/15237&quot;&gt;modes&lt;/a&gt; work coming from Jane Street to see how to support that. There&apos;s plenty more to do, so if you&apos;d like to lend a hand, reach out and join in!&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/04/odoc-3.html" rel="alternate" title="Odoc 3: So what?"/>
    <category term="odoc"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/04/semantic-versioning-is-hard.html</id>
    <title type="text">Semantic Versioning in OCaml is Hard</title>
    <updated>2025-04-20T00:00:00Z</updated>
    <published>2025-04-20T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;&lt;a href=&quot;https://semver.org&quot;&gt;Semantic versioning&lt;/a&gt; is a lovely and simple idea that, if it were reliably implemented everywhere, would make life a lot simpler. So, is it possible to make our OCaml libraries stick to this scheme? There are some projects that are trying to do this, including a recent &lt;a href=&quot;https://www.outreachy.org&quot;&gt;Outreachy&lt;/a&gt; project by &lt;a href=&quot;https://github.com/azzsal/&quot;&gt;Abdulaziz Alkurd&lt;/a&gt; mentored by &lt;a href=&quot;https://choum.net&quot;&gt;panglesd&lt;/a&gt; and &lt;a href=&quot;https://github.com/nathanreb&quot;&gt;Nathan Reb&lt;/a&gt;. While this is a great start, there are some subtleties of the OCaml module system that make it a good deal more complex than in other languages.&lt;/p&gt;
&lt;h2&gt;opam-format.2.3.0 ≠ opam-format.2.3.0?&lt;/h2&gt;
&lt;p&gt;Let&apos;s take the case that hit me this morning. I&apos;ve been working on &lt;a href=&quot;https://github.com/ocurrent/ocaml-docs-ci&quot;&gt;ocaml-docs-ci&lt;/a&gt; in order to bring the exciting new &lt;a href=&quot;https://ocaml.github.io/odoc&quot;&gt;odoc 3&lt;/a&gt; features to &lt;a href=&quot;https://ocaml.org/&quot;&gt;ocaml.org&lt;/a&gt; for everyone to enjoy. I have it checked out and building locally, but to deploy it to the infrastructure managed by &lt;a href=&quot;https://tunbury.org/&quot;&gt;Mark Elvers&lt;/a&gt; it needs to be packaged up into a Docker image. So I issued the usual &lt;code&gt;docker build .&lt;/code&gt; and after it churned through the setup stages and got on to building the project, it hit an error:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;File &amp;quot;src/solver/solver.ml&amp;quot;, line 58, characters 75-98:
     let deps = List.map (fun pkg -&amp;gt; OpamPackage.Map.find pkg simple_deps) (OpamPackage.Set.to_list pkgs) in
Error: Unbound value OpamPackage.Set.to_list
Hint: Did you mean of_list?
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now &lt;code&gt;OpamPackage&lt;/code&gt; is a module in the &lt;code&gt;opam-format&lt;/code&gt; library, which is easily discovered using the excellent &lt;a href=&quot;https://doc.sherlocode.com/?q=OpamPackage&quot;&gt;Sherlodoc&lt;/a&gt; tool, so I checked what version I had locally, and what version I had in the Docker container, and it turned out I was using exactly the same version -- 2.3.0 -- both locally and in the container. So what&apos;s going on?&lt;/p&gt;
&lt;p&gt;The problem is that the Dockerfile I was using was using OCaml version 4.14, whereas locally I was using OCaml 5.3. &amp;quot;But how on earth can this cause the API of &lt;code&gt;opam-format&lt;/code&gt; to change?&amp;quot; I hear you wail! Well, this is actually one of the simpler outcomes of the way the OCaml module system works. Let&apos;s look at &lt;a href=&quot;https://github.com/ocaml/opam/blob/2.3.0/src/format/opamPackage.mli&quot;&gt;the code&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The first thing we note is the absence of any definition of &lt;code&gt;Set&lt;/code&gt; or &lt;code&gt;Map&lt;/code&gt; here&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;where do they come from? It turns out they come from &lt;a href=&quot;https://github.com/ocaml/opam/blob/2.3.0/src/format/opamPackage.mli#L49&quot;&gt;this line here&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;include OpamStd.ABSTRACT with type t := t
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So let&apos;s take a look over in &lt;code&gt;opamStd.mli&lt;/code&gt; to see what that signature looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;(** A signature for handling abstract keys and collections thereof *)
module type ABSTRACT = sig

  type t

  val compare: t -&amp;gt; t -&amp;gt; int
  val equal: t -&amp;gt; t -&amp;gt; bool
  val of_string: string -&amp;gt; t
  val to_string: t -&amp;gt; string
  val to_json: t OpamJson.encoder
  val of_json: t OpamJson.decoder

  module Set: SET with type elt = t
  module Map: MAP with type key = t
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;OK, so we&apos;ve found the definitions of &lt;code&gt;Set&lt;/code&gt; and &lt;code&gt;Map&lt;/code&gt; - they refer to signatures &lt;code&gt;SET&lt;/code&gt; and &lt;code&gt;MAP&lt;/code&gt; which are defined just above in &lt;a href=&quot;https://github.com/ocaml/opam/blob/2.3.0/src/core/opamStd.mli#L17-L98&quot;&gt;opamStd.mli&lt;/a&gt;. Let&apos;s just look at &lt;code&gt;Set&lt;/code&gt; since that&apos;s where the problem was:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type SET = sig

  include Set.S

  val map: (elt -&amp;gt; elt) -&amp;gt; t -&amp;gt; t

  val is_singleton: t -&amp;gt; bool

  (** Returns one element, assuming the set is a singleton. Raises [Not_found]
      on an empty set, [Failure] on a non-singleton. *)
  val choose_one : t -&amp;gt; elt

  val choose_opt: t -&amp;gt; elt option

  val of_list: elt list -&amp;gt; t
  val to_list_map: (elt -&amp;gt; &apos;b) -&amp;gt; t -&amp;gt; &apos;b list
  val to_string: t -&amp;gt; string
  val to_json: t OpamJson.encoder
  val of_json: t OpamJson.decoder
  val find: (elt -&amp;gt; bool) -&amp;gt; t -&amp;gt; elt
  val find_opt: (elt -&amp;gt; bool) -&amp;gt; t -&amp;gt; elt option

  (** Raises Failure in case the element is already present *)
  val safe_add: elt -&amp;gt; t -&amp;gt; t

  (** Accumulates the resulting sets of a function of elements until a fixpoint
      is reached *)
  val fixpoint: (elt -&amp;gt; t) -&amp;gt; t -&amp;gt; t

  (** [map_reduce f op t] applies [f] to every element of [t] and combines the
      results using associative operator [op]. Raises [Invalid_argument] on an
      empty set, or returns [default] if it is defined. *)
  val map_reduce: ?default:&apos;a -&amp;gt; (elt -&amp;gt; &apos;a) -&amp;gt; (&apos;a -&amp;gt; &apos;a -&amp;gt; &apos;a) -&amp;gt; t -&amp;gt; &apos;a

  module Op : sig
    val (++): t -&amp;gt; t -&amp;gt; t (** Infix set union *)

    val (--): t -&amp;gt; t -&amp;gt; t (** Infix set difference *)

    val (%%): t -&amp;gt; t -&amp;gt; t (** Infix set intersection *)
  end

end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Sure enough, there&apos;s no &lt;code&gt;to_list&lt;/code&gt; function defined in there. Once again though, there&apos;s an &lt;code&gt;include Set.S&lt;/code&gt; in there. It turns out that that refers to the &lt;code&gt;Set&lt;/code&gt; module in the OCaml standard library. We can again &lt;a href=&quot;https://github.com/ocaml/ocaml/blob/5.3.0/stdlib/set.mli&quot;&gt;look at the source&lt;/a&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;val to_list : t -&amp;gt; elt list
    (** [to_list s] is {!elements}[ s].
        @since 5.1 *)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And there it is. The &lt;code&gt;to_list&lt;/code&gt; function has only been in the &lt;code&gt;Set&lt;/code&gt; module since version 5.1.&lt;/p&gt;
&lt;h2&gt;Using ocaml.org docs&lt;/h2&gt;
&lt;p&gt;It was pretty difficult to figure that out from the source, but happily there&apos;s a better way. We can browse the docs on https://ocaml.org/ - We can look at the docs for the &lt;a href=&quot;https://ocaml.org/p/opam-format/2.3.0/doc/OpamPackage/Set/index.html&quot;&gt;OpamPackage.Set module&lt;/a&gt; which, as of today, does not contain any &lt;code&gt;to_list&lt;/code&gt; function. The &lt;code&gt;include Set.S&lt;/code&gt; is there with the expansion showing all of the types and values coming from it, so we can click on the &lt;code&gt;Set.S&lt;/code&gt; link on the include line which takes us to the documentation for the stdlib from OCaml 4.11.2. Changing the version from the dropdown at the top to something more recent takes us to a page containing the &lt;code&gt;to_list&lt;/code&gt; function with the helpful &lt;code&gt;since 5.1&lt;/code&gt; annotation.&lt;/p&gt;
&lt;p&gt;This is, in fact, a relatively simple example of the sorts of issues that can occur that make semantic versioning a headache. In this example, it was a change due to a difference in the compiler version used, but there&apos;s nothing particularly special about that - a package may expose signatures derived from any of its dependencies! So is there anything we can do about this? Obviously, yes!&lt;/p&gt;
&lt;h2&gt;Towards a solution&lt;/h2&gt;
&lt;p&gt;Step 1 of any approach to solving this is to be able to identify which bits of a libraries API come from which packages, and therefore which versions of those packages. It turns out there may well be a nice way to piggy-back on a recent feature from Odoc, which was originally intended to help with suppressing suprious warnings.&lt;/p&gt;
&lt;p&gt;The problem we were tackling was that if your library ends up exporting a module whose signature is defined in someone else&apos;s package, then any warnings that come from it are unfixable. To fix this we added a tag to each signature of a module that indicates which package it originally came from. Odoc is then very careful to keep track of this as it performs its signature manipulations, resulting in an accurate way to know which signature elements came from which package. This fixed the problem of the spurious warnings nicely.&lt;/p&gt;
&lt;p&gt;Quite separately, we&apos;ve got the docs CI that is attempting to build docs for every version of every package. Obviously given the above, in order to exhaustively show all the possible APIs of every library, we should build all possible combinations of every version of every package. Clearly we can&apos;t possibly do this, so the docs CI focuses on the goal of building at least one solution for every version of every package.&lt;/p&gt;
&lt;p&gt;Now if you combine these two ideas, we can use the builds of the packages with the tracking of the package of the originating signatures to be able to precisely track the differences in API between different versions of a package. This would allow us to build a database of those changes, and with this in hand we can look at what APIs are used in any other package and be able to suggest upper and lower bounds on the versions of its dependencies.&lt;/p&gt;
&lt;p&gt;Now wouldn&apos;t that be cool?&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/04/semantic-versioning-is-hard.html" rel="alternate" title="Semantic Versioning in OCaml is Hard"/>
    <category term="ocaml"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/04/meeting-the-team.html</id>
    <title type="text">Meeting the Team</title>
    <updated>2025-04-08T00:00:00Z</updated>
    <published>2025-04-08T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;It&apos;s tremendously exciting to be back in the &lt;a href=&quot;https://www.cst.cam.ac.uk/&quot;&gt;Computer Laboratory&lt;/a&gt;, as the last time I worked here was just before the pandemic. I&apos;m now a member of the &lt;a href=&quot;https://www.cst.cam.ac.uk/research/eeg&quot;&gt;Energy and Environment Group&lt;/a&gt; whose goal is &amp;quot;to have a measurable impact on tools and techniques for de-risking the future&amp;quot;.&lt;/p&gt;
&lt;h2&gt;What&apos;s going on?&lt;/h2&gt;
&lt;p&gt;With such a broad goal, it&apos;s hard to know where to start and how I&apos;ll fit in, so my first few weeks have been spent getting to know the other members of the group and what they&apos;re up to. It&apos;s an incredibly inspiring group of individuals who are all doing amazing work, and it&apos;s really humbling and daunting to be a part of it.&lt;/p&gt;
&lt;p&gt;There&apos;s some really interesting work going on in our group on LLMs, principally led by the fantastic &lt;a href=&quot;https://toao.com/&quot;&gt;Sadiq Jaffer&lt;/a&gt;. We had a chat a few weeks ago and have started to explore some ideas around seeing how well LLMs can program in OCaml already before we start to do some RL training on them. Having not done any LLM stuff before, it&apos;s a steep learning curve for me, but we&apos;re already seeing some interesting results. We should have some more to say about this in the coming weeks.&lt;/p&gt;
&lt;p&gt;Last week I met with &lt;a href=&quot;https://digitalflapjack.com/&quot;&gt;Michael Dales&lt;/a&gt;, and he talked about the project &lt;a href=&quot;https://github.com/quantifyearth/shark&quot;&gt;shark&lt;/a&gt; that he and &lt;a href=&quot;patrick.sirref.org&quot;&gt;Patrick Ferris&lt;/a&gt; have been working on. It&apos;s kind of a mix between a shell and a jupyter-style notebook, with a strong focus on reproducibility. The traditional pain of notebooks is, of course, the execution model, whereby cells might be executed in any order you like. This means that the state you find the notebook in might not be even reachable again, let alone consistently reproducible. Shark is trying to address this by using file-system snapshots and clever analysis of the inputs and outputs of each cell to both ensure reproducibility, but also to allow a fast editing cycle, rerunning of only the bits that need to be rerun, even in the presence of slow data processing steps. It&apos;s a fascinating project, and I can&apos;t wait to see it in action when Michael gives us a demo!&lt;/p&gt;
&lt;p&gt;I also met up with &lt;a href=&quot;https://ryan.freumh.org&quot;&gt;Ryan Gibb&lt;/a&gt; with &lt;a href=&quot;https://www.dra27.uk/blog/&quot;&gt;David Allsopp&lt;/a&gt; and we had a good chat about his project &lt;a href=&quot;https://github.com/RyanGibb/babel&quot;&gt;Babel&lt;/a&gt;, which is using the &lt;a href=&quot;https://nex3.medium.com/pubgrub-2fb6470504f&quot;&gt;PubGrub&lt;/a&gt; algorithm to do package resolution for multiple package domains at once. We&apos;ve got a number of avenues to explore here, from building a PubGrub implementation in OxCaml, to using Babel to construct Docker images for opam packages entirely from scratch, without using a base image.&lt;/p&gt;
&lt;p&gt;With my other hat on as a member of the CTO office at &lt;a href=&quot;https://tarides.com/&quot;&gt;Tarides&lt;/a&gt;, I&apos;m very much looking forward to using OCaml and OxCaml to solve some real-world problems that are in an entirely different domain than I&apos;ve been used to over the last few years.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/04/meeting-the-team.html" rel="alternate" title="Meeting the Team"/>
    <category term="meta"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/04/this-site.html</id>
    <title type="text">This site</title>
    <updated>2025-04-07T00:00:00Z</updated>
    <published>2025-04-07T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;I&apos;ve spent a &lt;em&gt;lot&lt;/em&gt; of time over the past few years working on Odoc, the OCaml documentation generator, so when it came time to (re)start my own website and blog, I found it hard to resist thinking about how I might use odoc as part of it. We&apos;ve spent a lot of time recently trying to make odoc more able to generate structured documentation sites, so I&apos;ve gone all in and am trialling using it as a tool to generate my entire site. This is a bit of an experiment, and I don&apos;t know how well it will work out, but let&apos;s see how it goes.&lt;/p&gt;
&lt;p&gt;Additionally, I&apos;ve recently been working on a project currently called &lt;code&gt;odoc_notebook&lt;/code&gt;, which is a set of tools to allow odoc &lt;code&gt;mld&lt;/code&gt; files to be used as a sort of Jupyter-style notebook. The idea is that you can write both text and code in the same file, and then run the code in the notebook interactively. Since I&apos;ve only got a webserver, all the execution of code has to be done client side, so I&apos;m making extensive use of the phenomenal &lt;a href=&quot;https://github.com/ocsigen/js_of_ocaml&quot;&gt;Js_of_ocaml&lt;/a&gt; project to get an OCaml engine running in the browser.&lt;/p&gt;
&lt;p&gt;My focus has initially been on getting &apos;toplevel-style&apos; code execution working. As an example, let&apos;s write a little demo.&lt;/p&gt;
&lt;h2&gt;Demo&lt;/h2&gt;
&lt;p&gt;Let&apos;s start with a little demo:&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;let x = 1 + 2&lt;/x-ocaml&gt;
It&apos;s intended to look like an OCaml toplevel session, so each new expression starts with a &lt;code&gt;#&lt;/code&gt; and is terminated with a double semicolon. The response from the toplevel is then below that indented with 2 spaces. Right now, there&apos;s not much in the way of error checking so you can make it all very confused by deleting the hash, removing the &lt;code&gt;;;&lt;/code&gt; and so on. Avoiding this, however, you can edit the numbers here and hit &apos;run&apos; (maybe twice!) to see the results being updated.&lt;/p&gt;
&lt;p&gt;There is also a little integration to allow the code to produce output more interesting than just text. The following cell creates an SVG image and &apos;pushes&apos; it to &lt;code&gt;Mime_printer&lt;/code&gt;, which receives the mime value and renders it in the browser below the code block.&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;let svg = [
{|&amp;lt;svg height=&amp;quot;210&amp;quot; width=&amp;quot;500&amp;quot; xmlns=&amp;quot;http://www.w3.org/2000/svg&amp;quot;&amp;gt;|};
{|&amp;lt;polygon points=&amp;quot;100,10 40,198 190,78 10,78 160,198&amp;quot; |};
{|style=&amp;quot;fill:lime;stroke:purple;stroke-width:5;&amp;quot;/&amp;gt;&amp;lt;/svg&amp;gt;|}];;&lt;/p&gt;
&lt;p&gt;Mime_printer.push &amp;quot;image/svg&amp;quot; (String.concat &amp;quot;\n&amp;quot; svg)&lt;/x-ocaml&gt;&lt;/p&gt;
&lt;h2&gt;Things to come&lt;/h2&gt;
&lt;h3&gt;Merlin support&lt;/h3&gt;
&lt;p&gt;There are a bunch of things I want to add to this, for example, Merlin support. In fact, &lt;a href=&quot;https://github.com/voodoos/merlin-js&quot;&gt;merlin-js&lt;/a&gt; already exists and works, thanks to the fantastic work of &lt;a href=&quot;https://github.com/voodoos&quot;&gt;Ulysse&lt;/a&gt;, but the problem is that it&apos;s not really designed for toplevel work, and it doesn&apos;t work when the code is broken up into chunks like I do here. So either I need to concatenate all the cells together before I give it to Merlin, or I need to make each cell it&apos;s own little module and &apos;open&apos; every previous cell&apos;s module.&lt;/p&gt;
&lt;p&gt;Within a single cell, it does already work. You can see that Merlin is correctly underlining the error in the following cell. You should also be able to hover over the variables and see their types.&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;type t = { foo : int; bar : string };;&lt;/p&gt;
&lt;p&gt;let x = { foo = 1; bar = &amp;quot;hello&amp;quot; };;&lt;/p&gt;
&lt;p&gt;let this_line_has_an_error = { foo = 1; bar = None };;&lt;/x-ocaml&gt;
But across cells, I&apos;ve broken Merlin, though the code is executes correctly. You can see the problem in the following cell, which re-pushes the SVG image using the variable &lt;code&gt;svg&lt;/code&gt; defined in the cell above. Merlin highlights the use of the varible &lt;code&gt;svg&lt;/code&gt; is, because it&apos;s not aware of the varible, but the code gets executed correctly and the image is rendered below the cell.&lt;/p&gt;
&lt;p&gt;&lt;x-ocaml mode=&quot;interactive&quot;&gt;Mime_printer.push &amp;quot;image/svg&amp;quot; (String.concat &amp;quot;\n&amp;quot; svg)&lt;/x-ocaml&gt;
Edit 2025-05-20: I have now got merlin working across cells, though I&apos;m not convinced the current solution is the right long-term solution. S&lt;/p&gt;
&lt;h3&gt;Dynamic libraries&lt;/h3&gt;
&lt;p&gt;Currently the use of libraries it quite limited - they are defined more-or-less statically. I&apos;ve had dynamic libraries working in the past, but I need to re-implement them. The plan is to have the &lt;code&gt;cma&lt;/code&gt; files converted to &lt;code&gt;js&lt;/code&gt; files and then load them on-demand when the notebook specifies them. The tricky thing here is that we need to be able to use them both in the browser and in bytecode executables so that the &apos;test-promote&apos; workflow still works. This will probably require specifying the libraries by name, and having to re-implement the work that &lt;a href=&quot;https://projects.camlcity.org/projects/findlib.html&quot;&gt;findlib&lt;/a&gt; does to find the libraries and load them and their dependencies in the right order, though this time entirely over HTTP.&lt;/p&gt;
&lt;h3&gt;Other things&lt;/h3&gt;
&lt;p&gt;There are loads of other things I&apos;m interested in doing, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Investigating how to do &apos;exercises&apos; to allow readers to try things out in a guided way&lt;/li&gt;
&lt;li&gt;&apos;Test cells&apos; to see if implementations are correct&lt;/li&gt;
&lt;li&gt;Persistence of the notebook state - both using local and remote storage&lt;/li&gt;
&lt;li&gt;Integration of docs&lt;/li&gt;
&lt;li&gt;Exploration of the execution model - how to run the code in the right order and ensure reproducibility&lt;/li&gt;
&lt;li&gt;Use of remote execution engines rather than just in the browser&lt;/li&gt;
&lt;li&gt;Other languages?
Right now though, my focus is on the functionality required for this blog, with a secondary goal of looking at how we might use this sort of technology on the docs site on ocaml.org. Wouldn&apos;t it be cool to be able to drop into a live OCaml toplevel for any library in opam?&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Example notebooks&lt;/h2&gt;
&lt;p&gt;As a more extended example of odoc notebooks, I have converted to this format the course that I help teach at the University of Cambridge; &lt;a href=&quot;https://www.cl.cam.ac.uk/teaching/2425/FoundsCS/&quot;&gt;Foundations of Computer Science&lt;/a&gt;. &lt;a href=&quot;/notebooks/foundations/&quot;&gt;Try them out for yourself!&lt;/a&gt;.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/04/this-site.html" rel="alternate" title="This site"/>
    <category term="odoc"/>
    <category term="meta"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/03/module-type-of.html</id>
    <title type="text">The Road to Odoc 3: Module Type Of</title>
    <updated>2025-03-08T00:00:00Z</updated>
    <published>2025-03-08T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;There are &lt;a href=&quot;https://discuss.ocaml.org/t/ann-odoc-3-beta-release/16043&quot;&gt;many new and improved features&lt;/a&gt; that Odoc 3 brings, but there are also a large number of bugfixes. I thought I&apos;d write about one in particular here, an &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1081&quot;&gt;overhaul of &amp;quot;module type of&amp;quot;&lt;/a&gt; that landed in May 2024.&lt;/p&gt;
&lt;h2&gt;Module Type Of&lt;/h2&gt;
&lt;p&gt;module type of is a language feature of OCaml allowing one to recover the signature of an existing module. For example, if I had a module &lt;code&gt;X&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module X = struct
  type t = Foo | Bar
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;then I can get back the signature of &lt;code&gt;X&lt;/code&gt; using &lt;code&gt;module type of&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type Xsig = module type of X
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;which can be very useful if you’re trying to &lt;a href=&quot;https://discuss.ocaml.org/t/extend-existing-module/1389&quot;&gt;extend existing modules&lt;/a&gt; amongst other things.&lt;/p&gt;
&lt;p&gt;OCaml and Odoc treat module type of in somewhat different ways. OCaml internally expands the expression immediately it sees it, and effectively replaces it with the signature - ie, in the above example Xsig is now a signature, not a module type of expression.&lt;/p&gt;
&lt;p&gt;In contrast, Odoc would like to keep track of the fact that this signature came from a &lt;code&gt;module type of&lt;/code&gt; expression, as it’s very useful to know. If you’re extending a module, your signature might look like:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type UnitExtended = sig
  include module type of Unit
  val extra_unit_function : unit -&amp;gt; unit
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The documentation we produce will expand the contents of the &lt;code&gt;include&lt;/code&gt; statement, but keep track of the fact that it came from a &lt;code&gt;module type of&lt;/code&gt; expression so the reader can see where these signature items came from. In practice, you&apos;d probably want to use &lt;code&gt;module type of struct include Unit end&lt;/code&gt;, which is a bit different from simply &lt;code&gt;module type of Unit&lt;/code&gt;, and I&apos;ll talk about this at some point in a future post.&lt;/p&gt;
&lt;h2&gt;The problem&lt;/h2&gt;
&lt;p&gt;We run into difficulties as soon as we introduce another language feature that operates on signatures: with. Let’s start with a module type &lt;code&gt;S&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type S = sig
  module X : sig
    type t = int
  end

  module type Y =
    module type of X
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We’ll now define a new module &lt;code&gt;X2&lt;/code&gt; that we intend to use as a replacement for &lt;code&gt;X&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module X2 = struct
  type t = int
  type u = float
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now we’ll define a new module type &lt;code&gt;T&lt;/code&gt; which is &lt;code&gt;S&lt;/code&gt; but with &lt;code&gt;X&lt;/code&gt; replaced:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type T = S with module X := X2
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here you can see that OCaml has expanded the &lt;code&gt;module type of&lt;/code&gt; expressions and told us the computed signature. The interesting thing here is that in module type &lt;code&gt;T&lt;/code&gt;, module type &lt;code&gt;Y&lt;/code&gt; only has a type &lt;code&gt;t&lt;/code&gt; in it, not a type &lt;code&gt;u&lt;/code&gt;. As above, Odoc wants to keep the &lt;code&gt;module type of&lt;/code&gt; expression so the reader can tell where module type &lt;code&gt;Y&lt;/code&gt; came from. However, the substitution would do a different thing in this case - we would have the following:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-ocaml&quot;&gt;module type T = sig
  module type Y = module type of X2
end
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and the expansion of this would then clearly have both types &lt;code&gt;t&lt;/code&gt; and &lt;code&gt;u&lt;/code&gt; in it.&lt;/p&gt;
&lt;p&gt;So now Odoc has two problems: We need to compute the correct signature, and we need to be able to describe how we computed it.&lt;/p&gt;
&lt;h2&gt;The solution&lt;/h2&gt;
&lt;p&gt;The previous solution to this was to have a ‘phase 0’ of odoc which would compute the expansions of all module type of expressions before doing any other work. This was necessary because of a ‘simplfying’ assumption in how we handled the typing environment. The new, simpler approach was to calculate the expansion during the normal flow of work, and never to attempt to recalculate it, but simply operate on the signature. This was a nice big simplification and optimisation that removed a few corner cases in the previous code (including an &lt;a href=&quot;https://github.com/ocaml/odoc/blob/v2.4/src/xref2/type_of.ml#L167-L174&quot;&gt;infinite loop&lt;/a&gt; that we &lt;em&gt;hoped&lt;/em&gt; always terminated…!)&lt;/p&gt;
&lt;p&gt;The second issue was how to describe it. We still want it clear that this signature was derived from another, but it’s clear we can’t honestly say that in the above example that it’s &lt;code&gt;module type of X2&lt;/code&gt;. The answer is that we have applied a transparent ascription to the signature. Essentially, the signature is &lt;code&gt;X2&lt;/code&gt; but constrained to only have the fields of &lt;code&gt;X&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This is not a current feature of OCaml, though Jane Street has &lt;a href=&quot;https://blog.janestreet.com/plans-for-ocaml-408/&quot;&gt;done some work&lt;/a&gt; on this, including declaring the syntax: &lt;code&gt;X2 &amp;lt;: X&lt;/code&gt;. However, there’s another interesting wrinkle here. &lt;code&gt;X&lt;/code&gt; is a module defined in the module type &lt;code&gt;S&lt;/code&gt;, so it’s not possible to write a valid OCaml path that points to it – &lt;code&gt;S.X&lt;/code&gt; has no meaning. In addition, the right-hand side of the &lt;code&gt;&amp;lt;:&lt;/code&gt; operator should be a module type, so we’d actually need to write &lt;code&gt;X2 &amp;lt;: module type of S.X&lt;/code&gt; . We’re still figuring out the right thing to do here, so for now Odoc 3 will still pretend that it’s simply &lt;code&gt;module type of X2&lt;/code&gt;.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/03/module-type-of.html" rel="alternate" title="The Road to Odoc 3: Module Type Of"/>
    <category term="ocaml"/>
    <category term="odoc"/>
  </entry>
  <entry>
    <id>https://jon.recoil.org/blog/2025/03/code-block-metadata.html</id>
    <title type="text">Code block metadata</title>
    <updated>2025-03-07T00:00:00Z</updated>
    <published>2025-03-07T00:00:00Z</published>
    <content type="html">
      &lt;p&gt;Back in 2021 &lt;a href=&quot;https://github.com/julow&quot;&gt;julow&lt;/a&gt; introduced some &lt;a href=&quot;https://github.com/ocaml-doc/odoc-parser/pull/2&quot;&gt;new syntax&lt;/a&gt; to odoc’s code blocks to allow us to attach arbitrary metadata to the blocks. We imposed no structure on this; it was simply a block of text in between the language tag and the start of the code block. Now odoc needs to use it itself, we need to be a bit more precise about how it’s defined.&lt;/p&gt;
&lt;p&gt;The original concept looked like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{@ocaml metadata goes here in an unstructured way[
  ... code ...
]}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where everything in between the language (“ocaml” in this case) and the opening square bracket would be captured and put into the AST verbatim. Odoc itself has had no particular use for this, but it has been used in &lt;a href=&quot;https://github.com/realworldocaml/mdx&quot;&gt;mdx&lt;/a&gt; to control how it handles the code blocks, for example to skip processing of the block, to synchronise the block with another file, to disable testing the block on particular OSs and so on.&lt;/p&gt;
&lt;p&gt;As part of the Odoc 3 release we decided to address one of our &lt;a href=&quot;https://github.com/ocaml/odoc/pull/303&quot;&gt;oldest open issues&lt;/a&gt;, that of extracting code blocks from mli/mld files for inclusion into other files. This is similar to the file-sync facility in mdx but it works in the other direction: the canonical source is in the mld/mli file. In order to do this, we now need to use the metadata so we can select which code blocks to extract, and so we needed a more concrete specification of how the metadata should be parsed.&lt;/p&gt;
&lt;p&gt;We looked at what &lt;a href=&quot;https://github.com/realworldocaml/mdx/blob/main/lib/label.ml#L195-L210&quot;&gt;mdx does&lt;/a&gt;, but the way it works is rather ad-hoc, using very simple String.splits to chop up the metadata. This is OK for mdx as it’s fully in charge of what things the user might want to put into the metadata, but for a general parsing library like odoc.parser we need to be a bit more careful. Daniel Bünzli &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1326#issuecomment-2702260053&quot;&gt;suggested&lt;/a&gt; a simple strategy of atoms and bindings inspired by s-expressions. The idea is that we can have something like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{@ocaml atom1 &amp;quot;atom two&amp;quot; key1=value1 &amp;quot;key 2&amp;quot;=&amp;quot;value with spaces&amp;quot;[
    ... code content ...
]}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Daniel suggested a very minimal escaping rule, whereby a string could contain a literal &amp;quot; by prefixing with a backslash - something like; &amp;quot;value with a \&amp;quot; and spaces&amp;quot;, but we discussed it during the &lt;a href=&quot;https://ocaml.org/governance/platform&quot;&gt;odoc developer meeting&lt;/a&gt; and felt that we might want something a little more familiar. So we took a look at the lexer in &lt;a href=&quot;https://github.com/janestreet/sexplib/blob/master/src/lexer.mll&quot;&gt;sexplib&lt;/a&gt; and found that it follows the &lt;a href=&quot;https://github.com/janestreet/sexplib/blob/d7c5e3adc16fcf0435220c3cd44bb695775020c1/README.org#lexical-conventions-of-s-expression&quot;&gt;lexical conventions&lt;/a&gt; of OCaml’s strings, and decided that would be a reasonable approach for us to follow too.&lt;/p&gt;
&lt;p&gt;The resulting code, including the extraction logic, was implemented in &lt;a href=&quot;https://github.com/ocaml/odoc/pull/1326/&quot;&gt;PR 1326&lt;/a&gt; mainly by &lt;a href=&quot;https://github.com/panglesd&quot;&gt;panglesd&lt;/a&gt; with a little help from me on the lexer.&lt;/p&gt;

    </content>
    <link href="https://jon.recoil.org/blog/2025/03/code-block-metadata.html" rel="alternate" title="Code block metadata"/>
    <category term="odoc"/>
  </entry>
</feed>