From 3c03c9b03a084984383b510f6b795df3bfaf5ca7 Mon Sep 17 00:00:00 2001 From: muqiuhan Date: Fri, 18 Oct 2024 15:45:18 +0000 Subject: deploy: e8ad1ffc5e1b38c1ac5bb25141187398e15d1236 --- .../index.html | 425 +++++++++++++++++++ .../index.html | 463 +++++++++++++++++++++ .../index.html | 370 ++++++++++++++++ 2024/10/18/Generalised-signature/index.html | 383 +++++++++++++++++ 4 files changed, 1641 insertions(+) create mode 100644 2024/10/18/10-Tips-for-Productive-FSharp-Scripting/index.html create mode 100644 2024/10/18/Building-custom-fibers-library-in-FSharp/index.html create mode 100644 2024/10/18/Dealing-with-complex-dependency-injection-in-FSharp/index.html create mode 100644 2024/10/18/Generalised-signature/index.html (limited to '2024/10/18') diff --git a/2024/10/18/10-Tips-for-Productive-FSharp-Scripting/index.html b/2024/10/18/10-Tips-for-Productive-FSharp-Scripting/index.html new file mode 100644 index 00000000..efff6f07 --- /dev/null +++ b/2024/10/18/10-Tips-for-Productive-FSharp-Scripting/index.html @@ -0,0 +1,425 @@ + + + + + + + + + + + + + + + + + + + + + +10 Tips for Productive FSharp Scripting | 暮秋小屋 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+
+ +
+ +
+
+
+ + + +
+
+
+ + +
+
+
+ + +
+ +
+ +
+ +
+
+
+

Scott Hanselman recently had a nice post on C# and F# REPLs, which reminded me of the time I started using F# scripts. Over time, I found out a couple of small tricks, which helped make the experience productive. I found about them mainly by accident, so I figured, let’s see if I can list them in one place! Some of these are super simple, some probably a bit obscure, but hopefully, one of them at least will make your path towards scripting nirvana an easier one…

+
+

Note: these tips are not necessarily ordered by usefulness. For that matter, there might or might not be exactly 10 of them :)

+
+

Tip 1: Use .fsx Files for Interactive Coding

You can use the F# Interactive 2 ways: you can directly type code into FSI, the F# Interactive window, or you can write code in an .fsx file, and select pieces of the code you want to execute. I recommend the second approach, for at least two reasons. First, FSI is a very primitive environment, .fsx files provide a much richer experience (IntelliSense). Then this encourages writing clean scripts you can reuse later.

+
+

This is not specific to scripts, but… if you are on Visual Studio, do yourself a service and install the Visual F# Power Tools - you’ll get nice things such as better code highlighting, refactoring, and more.

+
+

To execute code interactively, simply type code in an .fsx file, select a block of code, and hit Alt + Enter. The selected code will be evaluated, and the result will show up in the FSI window. In Visual Studio, you can also select code and right-click “Execute in Interactive”, but shortcuts are way faster.

+
+

You can also execute a single-line with Alt + ‘. I rarely use this option, but this can save you time because you don’t need to select the entire line of code.

+
+
+

In case the keyboard shortcuts to send code to FSI do not work anymore (ReSharper used to over-write them in the past), you can reset them in Visual Studio, by going to Tools / Options / Environment / Keyboard. The 2 commands you need to map are EditorContextMenus.CodeWindow.ExecuteInInteractive and EditorContextMenus.CodeWindow.ExecuteLineInInteractive.

+
+

You can also use these shortcuts from a regular .fs file, which can be handy if you want to validate that a piece of code is behaving the way you want.

+
+

Interactive coding is by far my main usage for scripts - I use it extensively to prototype designs, run dumb tasks, or explore data or libraries. I realized recently that a few of my C# friends use LinqPad for the same purpose.

+
+

Tip 2: What is it?

While I encourage working primarily from .fsx files, the FSI window is also very helpful. I use it primarily for small verifications. For instance, I might have in my script file code like this:

+
1
2
let add x y =
x + y
+ +

Once I send it for evaluation into FSI, I will see the following show up in FSI:

+
1
2
val add : x:int -> y:int -> int
>
+ +

My function add is now in memory, in my FSI session; I can start typing in the FSI window and use it:

+
1
2
3
> add 1 2;;
val it : int = 3
>
+ +

Enter does not trigger execution in FSI. The ;; indicates to FSI “Please execute everything I just typed, up to that point”. This is useful if you want to type multiple lines of code in FSI, and execute them as a block.

+
+

it: in our add 1 2 example, the result showed up as it. We simply ran add, but didn’t assign the result to anything. it now contains the result, until we run another expression. If you want to re-use that value, you can assign it in FSI, by doing for instance let x = it;;.′

+
+
+

Once a value is loaded in your FSI session, it will remain there, available to you until you shadow it (in the example above, x will remain available, until I run for instance let x = 42;;). This is extremely convenient: for instance, you can load a data file once let data = File.ReadAllLines path, and keep using data for as long as you want, without having to reload it between code changes.

+
+
+

FSI often shows an abbreviated version of values for large items. For instance, [1..999] will show up as val it : int list = [1; 2; 3; 4; 5; 6; 7; 8; 9; 10; 11; 12; 13; 14; 15; 16; 17; 18; 19; 20; 21; 22; 23; 24; 25; 26; 27; 28; 29; 30; 31; 32; 33; 34; 35; 36; 37; 38; 39; 40; 41; 42; 43; 44; 45; 46; 47; 48; 49; 50; 51; 52; 53; 54; 55; 56; 57; 58; 59; 60; 61; 62; 63; 64; 65; 66; 67; 68; 69; 70; 71; 72; 73; 74; 75; 76; 77; 78; 79; 80; 81; 82; 83; 84; 85; 86; 87; 88; 89; 90; 91; 92; 93; 94; 95; 96; 97; 98; 99; 100; ...] - note the … at the end, which indicate that there is more.

+
+

What if you inadvertently started a very long computation, or an infinite loop? In Visual Studio, you can either kill the session entirely, by right-clicking over the FSI window and selecting “Reset Interactive Session” or Ctrl + Alt + R, or cancel the latest evaluation you requested (“Cancel Interactive Evaluation”, or Ctrl + Break.).

+

Tip 3: Run Scripts from the Command Line

Besides interactive scripting, you can also run a script from the command line, by using FSI.exe:

+

>fsi.exe "C:\myscript.fsx"

+
+

FSI.exe is typically located at C:\Program Files (x86)\Microsoft SDKs\F#\4.0\Framework\v4.0. You can also install it separately, see fsharp.org/use section for instructions for various platforms.

+
+

You can define different behaviors in your script, depending on whether it is run interactively or from the command line, like this:

+
1
2
3
4
5
6
7
#if INTERACTIVE
let msg = "Interactive"
#else
let msg = "Not Interactive"
#endif

printfn "%s" msg
+ +

Updated, Sep 19: thanks Matt Klein for pointing the issue.

+

For more information on FSI from the command line, check the reference page here.

+

Updated, Feb 20: Ramon Soto Mathiesen points out that Tip 9 also applies to the command line.

+

Tip 4: Use Relative Paths

Sometimes, your script will reference another resource; for instance, you need to read the contents of a .txt file somewhere. You can use absolute path, as in:

+
1
File.ReadAllLines @"C:/data/myfile.txt"
+ +
+

Pre-pending a string with @ makes it a verbatim string, and ignore escape sequences, such as \.

+
+
+

Use / rather than \, so that path work both on Windows and Mono.

+
+

However, if that resource lives in a location relative to your script, consider using relative path, so that you can move your script folder around without breaking it.

+

Relative paths can be a bit tricky; for instance, running the following code interactively…

+
1
System.Environment.CurrentDirectory
+ +

… produces a potentially unexpected result in FSI:

+
1
2
val it : string = "C:\Users\Mathias Brandewinder\AppData\Local\Temp"
>
+ +

You can avoid these issues by using built-in constants, which refer respectively to the directory where the script lives, the script file name, and the current line of the script:

+
1
2
3
__SOURCE_DIRECTORY__
__SOURCE_FILE__
__LINE__
+ +

So if your folder structure was along these lines…

+
1
2
3
root
/src/script.fsx
/data/data.txt
+ +

… you could refer to the data file data.txt from your script like this:

+
1
2
let path = System.IO.Path.Combine(__SOURCE_DIRECTORY__,"..","data/data.txt")
System.IO.File.ReadAllText path
+ +

Tip 5: Including Assemblies

By default, FSI loads FSharp.Core and nothing else. If you want to use System.DateTime, you will need to first open System in your script. If you want to use an assembly that is not part of the standard .NET distribution, you will need to reference it first using #r. Imagine for instance that you installed the Nuget package fsharp.data; to use it in your script, you would do something like:

+
1
2
#r @"../packages/FSharp.Data.2.2.5/lib/net40/FSharp.Data.dll"
open FSharp.Data
+ +
+

When you execute open System in interactive, don’t worry if nothing seems to happen: the only result is a new > showing up in FSI.

+
+

For assemblies that are part of .NET but not referenced by default, you can use a shorter version:

+
1
2
#r @"System.Xaml"
open System.Xaml
+ +
+

In Visual Studio, you can right-click a reference from Solution Explorer, and send to F# interactive. You can then directly open it, and start using it in FSI.

+
+

Updated, Feb 20: Sergey Tihon shared an interesting comment, explaining where Tip 5 can sometimes go wrong. I’d say, try Tip 5 first, but be aware that this might at times not quite work:

+
+

@brandewinder don’t load assemblies like in Tip 5 ) https://t.co/Owft1NmPoo

+

— Sergey Tihon (@sergey_tihon) February 7, 2016

+
+

Updated, Feb 20: F# open source contributor Don Syme share a related nice trick:

+
+

@jeroldhaas @sergey_tihon @brandewinder Use #I SOURCE_DIRECTORY, it is wondrous, very satisfying. All relative paths then work

+

— Don Syme (@dsyme) February 7, 2016

+
+

Tip 6: Use Paket

The Nuget package manager is useful to consume existing packages. However, by default, Nuget stores assemblies in a folder that includes the package version number. This is very impractical for a script. In our example above, if fsharp.data gets an update, our script reference will be broken once we update the Nuget package:

+

#r @"../packages/FSharp.Data.2.2.5/lib/net40/FSharp.Data.dll"

+

Fixing the script requires manually editing the version number in the path, which quickly becomes a pain. Paket provides a better experience, because it stores packages without the version number, in this case, under:

+

#r @"../packages/FSharp.Data/lib/net40/FSharp.Data.dll"

+

Your scripts will now gracefully handle version number changes.

+

If you end up consuming numerous packages, you can make your life even easier, by referencing paths where assemblies might be searched for, using #I:

+
1
2
#I @"../packages/
#r @"FSharp.Data/lib/net40/FSharp.Data.dll"
+ +
+

If your primary goal is to “just script”, consider using Atom or VSCode, with the Ionide plugin. You can create and run free-standing F# scripts, with beautiful Paket integration.

+
+

Tip 7: Include Files

You might want to use the code from an existing file in your script. Suppose that we have a code file Code.fs somewhere, looking like this:

+
1
2
3
4
namespace Mathias

module Common =
let hello name = sprintf "Hello, %s" name
+ +

You can use that code from your script, by using the #load directive:

+
1
2
3
#load "Code.fs"
open Mathias.Common
hello "World"
+ +
+

You might have to close and re-open the script file if you end up changing the contents of the file.

+
+
+

If the file you are attempting to load contains references to other assemblies or files, you might get an error on the #load statement: “One or more errors in loaded files. The namespace or module … is not defined”. Simply reference the missing assemblies above the #load statement, so that your script uses the same dependencies as the file it refers to.

+
+

Tip 8: Profile your Code with #time

Another handy directive, #time, turns on basic profiling. Once it is executed, for every block of code you send for execution you will see timing and garbage collection information. For instance, running this code…

+
1
2
#time
[| 1 .. 10000000 |] |> Array.map (fun x -> x * x)
+ +

… will produce the following in FSI:

+
1
2
3
4
5
--> Timing now on

Real: 00:00:00.887, CPU: 00:00:00.828, GC gen0: 2, gen1: 2, gen2: 2
val it : int [] =
[|1; 4; 9; 16; 25; 36; 49; // snipped for brevity
+ +

We get the wall time and CPU time it took, as well as some information about garbage collection in generations 0, 1 and 2. This would not replace a full-blown profiler, but this is an awfully convenient tool to figure out quickly if there are obvious ways to improve a piece of code.

+

Note that every time you execute #time, the timer will be switched from on to off, or vice-versa. This is not always convenient; you can also explicitly set it to the desired state, like this:

+
1
2
3
#time "on"
// everything now is timed
#time "off"
+ +
+

If you are interested in profiling, you should take a look at PrivateEye; check out Greg Young’s talk at NDC Oslo 2015 to get a feel for what it does.

+
+

Tip 9: Turn 64-bits on

Hat tip to Rick Minerich for that one. I’ll refer you to his blog post to see how to set FSI to 64 bits to handle large datasets.

+

Tip 10: Bonus Material

Did you know that you could…

+ +

And again… if you are not using the Visual F# Power Tools, you are missing out:

+
+

“Don’t let your friends try #fsharp without installing @FSPowerTools.” @dsyme at #ndclondon

+

— Tomas Petricek (@tomaspetricek) January 15, 2016

+
+

That’s what I got! I am sure I forgot some - do you have a useful or favorite trick to share?

+ +
+ + + + + +
+ + + + + + + +
+ + +
+
+
+ + + +
+ + +
+
+
+
+
+
+ +
+
+
+
+
+
+
+
+
+ + + + + + + + + diff --git a/2024/10/18/Building-custom-fibers-library-in-FSharp/index.html b/2024/10/18/Building-custom-fibers-library-in-FSharp/index.html new file mode 100644 index 00000000..13709c45 --- /dev/null +++ b/2024/10/18/Building-custom-fibers-library-in-FSharp/index.html @@ -0,0 +1,463 @@ + + + + + + + + + + + + + + + + + + + + + +Building custom fibers library in FSharp | 暮秋小屋 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+
+ +
+ +
+
+
+ + + +
+
+
+ + +
+
+
+ + +
+ +
+ +
+ +
+
+
+

Over the course of last few months on this blog post, I’ve been sharing about internals and how-to of different concurrency patters. We discussed how to implement our own actors and specific affinity-based thread pool. Today we’ll focus of the most dominant pattern present in modern programming nowadays: fibers, also known as coroutines, futures, tasks, green threads or user-space threads.

+

The general idea is simple - we want a fine-grained concurrency primitive, that will let us easily compose chain of operations in sequential manner. Of course we could use threads here, but the question is: are threads fine-grained? In many managed languages with OS threads exposed, they can be quite heavy eg. by default in .NET each thread takes around 1MB of memory and requires calling kernel code to cooperate with other threads, which is an expensive operation on its own.

+

What we’re after, are more lightweight structures (less than 1kB), that can live fully in a user space, so that we can have even millions of them cooperating frequently with each other without heavy performance penalties.

+

Before we begin, I think it’s good to discuss different designs. We’ll cover several different topics to be able to make more informed decisions, that we’re up to apply to our own solution.

+

Preemptive vs cooperative scheduler

Scheduler is a subsystem, which direct responsibility is to assign CPU core processing power to a particular fiber. It’s also responsible for coordinating fibers execution. The two most common categories of schedulers are preemptive and cooperative.

+

A preemptive scheduler is the one, that’s always in control of fiber execution. It’s able to decide on its own, when fiber can be started and stopped. The most obvious example of such is a thread scheduler existing on most operating systems.

+

Preemptive scheduler usually works in one of two ways:

+
    +
  • Time based scheduler takes a quant of CPU time and gives it to a given fiber, which ten can execute its logic until it reaches its execution time limit (of course, it can finish earlier). This is how OS thread scheduler, but also how Go goroutine scheduler works.
  • +
  • Another variant is step-based scheduler, which splits fiber’s function body into series of (more or less equal) steps. Then each fiber is given a number of steps to execute before preemption occurs. Example of such is Erlang’s BEAM - it simply allows each process to execute up to 2000 “reductions”, where each reduction is basically a function call. And since in Erlang there are no loops, only tail-recursive functions, this approach works well for long-living iterative processes as well.
  • +
+

One of the problems with preemptive schedulers is that they usually need some kind of involvement from the compiler or hosting virtual machine in order to work. For this reason, most of the fiber libraries use cooperative schedulers to perform their work.

+

A cooperative scheduler doesn’t have a concept of preemption - once started by the scheduler, a fiber will execute until it doesn’t give back the control willingly. This is often done with dedicated programming constructs, and often is known as yielding, parking or awaiting.

+

In cooperative variant, a fiber body is usually split into series of discrete steps, between which fiber gives control back to the scheduler.

+

Keep in mind that these two are not mutually exclusive - a preemptive scheduler often provides a way for a fiber to return control back to it when it’s known that fiber won’t be executing any longer eg. because it has been put to sleep for a while.

+

Stackless vs. stackful

A concept, that’s somewhat related to a topic above is the idea of stackless and stackful coroutines.

+

A stackful variant is aware of underlying execution stack and can preserve/restore parts of it when yielding/continuing a fiber. Examples of this approach could be Go, Lua, Python asyncio and in the future, also Java Loom project. Implementing such option (if it’s not implemented by a runtime already) usually requires diving deep into low-level internals, since execution stack is not something that most managed languages offers the users to play with, and doing so without coordination with runtime can cause problems - like determining liveness of objects for GC purposes.

+

Stackless coroutine usually captures locals that we want to preserve as part of callback object (lambda), that is allocated as an object on the heap and scheduled on yield continuation. These steps are usually visible directly in code (eg. await in C#, Rust and JavaScript, but also joints of Scala for-comprehensions, bang-suffix in F# or Haskell do-notation), but sometimes can be implicit like in case of Kotlin. Take into account that while many languages offer syntax support for those constructs, it’s not explicitly necessary to work - take a look at JavaScript and Promise.then as an example.

+

Stackless coroutines usually construct their logic around one of two concepts:

+
    +
  1. Finite state machines - this variant is usually faster and can be encoded manually (example of such case is Akka actors), but for a human eye it usually doesn’t really read as a sequential step-by-step program execution, unless it has some support from the compiler itself (see: C# and Rust).
  2. +
  3. Monadic sequencing via bind/flatMap operator, which is very popular in functional languages. While we cover it in more details in the rest of this blog post, for now it’s enough to say that it’s a way to chain callback-based behaviors together in a way, that resembles standard sequential code.
  4. +
+

For sure one of the advantages of stackful coroutines is that they’re mono-colored: you can yield/continue coroutine execution from within any other function, while in the stackless variant splits your world into two-colored functions - synchronous and asynchronous - where async one can be only called and yielded safely (without blocking underlying OS thread) from within another async function.

+

Eager vs lazy fibers

We already mentioned two important events in fiber execution life cycle - starting and parking. Here I briefly discuss about different design decisions on when to start a fiber execution.

+

Eager execution means, that fiber is started automatically after its creation. An example of such are Scala Future[A] and JavaScript Promise. Since execution process starts right away, we’re willingly resign from a certain degree of control over how or when to execute given fiber. Usually this is solved by wrapping a fiber creation into another function or lambda.

+

Lazy execution is much more common and preferred way of work, as it allows us to separate place where we want to define our asynchronous sequence of steps from the place, where the execution details are defined. It’s used in C# TPL as well as pretty much in all functional languages implementations (excluding Scala futures mentioned earlier).

+

Interruption

There are also few decisions regarding premature escaping the fiber execution, also known as interruption/cancelation: one of them requires passing special object - a token - between method calls and explicit checking for its completion. It is how C# Tasks work. However putting such requirement onto the API user can be cumbersome and error-prone option. Therefore pretty much every other coroutine library either allows to direcly interrupt a fiber or (like in case of F# Async) passes cancelation tokens and check if they were triggered under the hood.

+

Implementation

Since we talked a bit about various approaches, let’s get to the meat of this blog post: implementing our own coroutine library in F#. So, what properties will it have?:

+
    +
  1. We use cooperative scheduling (we don’t want to tweak the compiler) of stackless fibers with support from F# computation expression for nice syntax.
  2. +
  3. We use simple approach by defining custom bind operator with support from F# computation expressions. No state machines.
  4. +
  5. We’ll use lazy invocation.
  6. +
  7. We’ll make use of implicitly passed cancelation tokens. We’ll handle them directly inside the linking code.
  8. +
+

All of these give us in very similar approach to that found inside of native F# Async data type. To begin with, we’ll simply define the shape of our fiber.

+

Underneath, pretty much every cooperative stackless coroutine approach uses callbacks to drive the flow of synchronous segments of code to be executed one after another. So what we need is a callback which takes a result of previous coroutine and schedules in within some context of execution:

+
1
type Fiber<'a> = Fiber of (ExecutionContext -> FiberCallback<'a> -> unit)
+ +

Here we’ll represent Fiber as a simple single-case discriminated union. We could as well define other specialized cases, like:

+
    +
  • Situation when coroutine is executed immediately and doesn’t need to be awaited on: think about variant of ValueTask from C# Task Parallel Library.
  • +
  • Case when coroutine fails - in that case we might want to store an artificial tracing context that would allow us to create nicely-formatted “stack traces”: since .NET Core 2.1, C# already provides similar solution however AFAIK it’s been solved differently.
  • +
+

Ok, but what are ExecutionContext and FiberCallback<'a>? Let’s start from callback. We can represent it as follows:

+
1
type FiberCallback<'a> = FiberResult<'a> -> unit
+ +

It’s just a simple function, which takes result of previous fiber execution and handles it. What’s the FiberResult<'a> then?

+

Our fiber can complete successfully (returning a value) or fail with an exception. We’ll be conservative here and won’t go into more typed world of IO bifunctor. We can easy define these possible outputs in F# using Result<'a, exn>.

+

Question is: is that exhaustive? Well… no. As we already mentioned, there’s a 3rd state, often overlooked or conflated with failure: a canceled fiber. A canceled fiber doesn’t produce any output - since it was canceled before completion. In F# we already know how to represent an absence of value - simply use an option. Therefore our ultimate Fiber result type could look like this:

+
1
type FiberResult<'a> = Result<'a, exn> option
+ +

Now, the ExecutionContext. While it can be compound of many different capabilities throughout the system - even to serve as functional equivalent of dependency injection - here I’ll use it only for implicit passing of specific scheduler info and cancelation tokens from one fiber to another.

+
1
type ExecutionContext = IScheduler * Cancel
+ +

IScheduler interface is used to abstract component responsible for running our fibers. At the moment all we need is an ability to schedule fiber execution:

+
1
2
3
[<Interface>]
type IScheduler =
abstract Schedule: (unit -> unit) -> unit
+ +

While the name and signature imply multithreaded execution model, it doesn’t have to be the case. We can even implement scheduler which will simulate everything on a single core.

+

For now, we can simply implement a scheduler API on top of our standard .NET thread pool:

+
1
2
3
4
5
6
7
8
module Scheduler

open System.Threading

let shared =
{ new ISchedule with
member __.Schedule fn =
ThreadPool.QueueUserWorkItem(WaitCallback (fun _ -> fn())) |> ignore }
+ +

Cancellation

Now it’s a time for cancellation tokens. Of course we could just make use of a flag - conceptually working like native .NET CancellationToken. However given implicit cancellation, it may not be enough. Example:

+
+

Imagine, that inside our fiber we’re scheduling the race between two other fibers ie. one writing data to a file and other which will complete after timeout. Now, whenever one of them completes first, we want to cancel another one to stop wasting resources for result that no longer matters.

+
+

This simple scenario is similar to what .NET Task.WhenAny is used - with a difference that, unlike TPL, we want to actually cancel other executing tasks instead of letting them run (potentially forever) :D

+

Now, since our cancellation is not explicit, we need to deal with few things:

+
    +
  1. Whenever parent fiber is cancelled, all child fibers it spawned are also cancelled.
  2. +
  3. Whenever we cancel a fiber that loose the race, we don’t want to accidentally cancel a token of its parent.
  4. +
+

This behavior implies at least using two separate tokens, however in practice it will be more pragmatic to make our Cancel token work as a tree hierarchy - this way we can easily keep track of things and support more complex scenarios.

+
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
[<Sealed;AllowNullLiteral>]
type Cancel(parent: Cancel) =
let mutable flag: int = 0
let mutable children: Cancel list = []
new() = Cancel(null)
/// Check if token was cancelled
member __.Cancelled = flag = 1
/// Remove child token
member private __.RemoveChild(child) =
let rec loop child =
let children' = children
let nval = children' |> List.filter ((<>) child)
if not (obj.ReferenceEquals(children', Interlocked.CompareExchange(&children, nval, children')))
then loop child
if not (List.isEmpty children) then loop child
/// Create a new child token and return it.
member this.AddChild () =
let rec loop child =
let children' = children
if (obj.ReferenceEquals(children', Interlocked.CompareExchange(&children, child::children', children')))
then child
else loop child
loop (Cancel this)
/// Cancel a token
member this.Cancel() =
if Interlocked.Exchange(&flag, 1) = 0 then
for child in Interlocked.Exchange(&children, []) do child.Cancel()
if not (isNull parent) then parent.RemoveChild(this)

+ +

The general idea is simple: every new cancellation token (except root) may have a parent and a list of children. Canceling parent means canceling its children as well. After cancellation, we need to unpin child from its parent (therefore need for RemveChild operation) to avoid memory leaks.

+
Lock-free updates

What might be confusing for some in the code above, are recursive loops inside of AddChild/RemoveChild operations. This is a good place to introduce lock-free algorithms: we use atomic operations from Interlocked class to make sure that we can replace field references within a single CPU instruction, therefore making such field update safe without synchronized access. This is also known as Compare-And-Swap semantics.

+

This alone however is not enough, as Interlocked.CompareExchange(&field, new', old) can only safely replace a single field with new value if it contained an old one. This means that you cannot safely add or remove element to the list. So what can we do?

+
    +
  1. We’re taking a value from the field.
  2. +
  3. Update that value.
  4. +
  5. Conditionally put it back again. What if in the meantime the field was already replaced by another concurrently running thread? In that case Interlocked.CompareExchange will return field value other that the one we read in step 1. This is why we compare its result with the variable we expected.
  6. +
  7. If the expectation fails, we’ll retry - hence a recursive loop. Eventually even in high contention scenarios we should be able to complete after few retries. Given cheap and idempotent update operation, this still will be way faster than trying to call kernel code to obtain mutex/semaphore lock.
  8. +
+

While this may sound like something error prone - we can potentially add the same element multiple times - in practice it’s safe, because our collection here is an immutable data structure. Adding the same element multiple times without updating the reference will always produce the same result.

+

Back on track…

Now we have pretty much all core structures. We’re ready to start building our fiber operators. Starting from the basic ones - a successfully completed fiber and the failed one:

+
1
2
3
4
5
let success r = Fiber <| fun (_, c) next -> 
if c.Cancelled then next None else next (Some (Ok r))

let fail ex = Fiber <| fun (_, c) next ->
if c.Cancelled then next None else next (Some (Error ex))
+ +

Here, we simply pass a result/error to our Fiber callback:

+
    +
  • Cancelled fiber call next callback with None - as fibers cancelled before completion produce output .
  • +
  • Successful call results in passing Some (Ok result) to a callback…
  • +
  • … while failed result can be identified with Some (Error exception).
  • +
+

You’ll be able to see a cancellation check made here as preamble of pretty much every operator body, which we’ll define. While it may sound cumbersome remember: we do that so that users of our fibers won’t have to :)

+

Next very important operation is result mapping - we want to map result of one fiber into something else, returning another (lazy) fiber:

+
1
2
3
4
5
6
7
8
let mapResult (fn: Result<'a> -> Result<'b>) (Fiber call) = Fiber <| fun (s, c) next ->
if c.Cancelled then next None
else
try
call (s, c) (fun result ->
if c.Cancelled then next None
else next (Option.map fn result))
with e -> next (Some (Error e))
+ +

We can use this function to compose more traditionally-looking map function…

+
1
let map (fn: 'a -> 'b) fiber = mapResult (Result.map fn) fiber
+ +

… however mapResult is more powerful - you could easily imagine using to apply failure recovery (a.k.a try/catch semantics) by simply mapping Error exception → Ok recoveredValue:

+
1
let catch fn fiber = mapResult (function Error e -> fn e | ok -> ok) fiber
+ +

Another must-have function is binding operator (also know as flatMap in other languages like Scala, or Promise.then in JavaScript). It gives us the ability to compose fibers together - we’ll also use it when we come up to building a computation expression for our fibers.

+
1
2
3
4
5
6
7
8
9
10
11
12
13
let bind (fn: 'a -> Fiber<'b>) (Fiber call) = Fiber <| fun (s, c) next ->
if c.Cancelled then next None
else
try
call (s, c) (fun result ->
if c.Cancelled then next None
else match result with
| Some (Ok r) ->
let (Fiber call2) = fn r
call2 (s, c) next // pass `next` callback over to next fiber
| None -> next None
| Some (Error e) -> next (Some(Error e)))
with e -> next (Some(Error e))
+ +

It’s simple - we execute one fiber from within another, passing the next callback from outer function as an argument to inner one.

+

Fiber computation expressions

With these few functions we’re already prepared to build a basic computation expression, that will enable us programming with fibers in pleasant way:

+
1
2
3
4
5
6
7
8
9
10
11
[<Struct>]
type FiberBuilder =
member inline __.Zero = Fiber.success (Unchecked.defaultof<_>)
member inline __.ReturnFrom fib = fib
member inline __.Return value = Fiber.success value
member inline __.Bind(fib, fn) = Fiber.bind fn fib

[<AutoOpen>]
module FiberBuilder =

let fib = FiberBuilder()
+ +

While in F# there are many more operators we could pack into our computation expression, these are basic ones that will let it work. With such construct, we’ll be able to write programs like:

+
1
2
3
4
5
6
7
8
9
let inline millis n = TimeSpan.FromMilliseconds (float n)

let program: Fiber<int> = fib {
let a = fib {
do! Fiber.delay (millis 1000) // create some artificial delay
return 3
}
let! b = a |> Fiber.timeout (millis 3000) // execute task within specified timeout
return b }
+ +

Sure, we have neither delay nor timeout operators at the moment, but at least you know where are we heading now :)

+

Delayed execution

In order to implement delays, we could theoretically just call Thread.Sleep and get over it, but this approach is devastating from any coroutine library point of view. Most user-space thread libraries work by using a predefined fixed pool of OS-level threads and scheduling coroutines on them - you can read more about building thread pools here.

+

However, Thread.Sleep(timeout) doesn’t know thread pooling mechanism - all it knows about is that we called suspending current OS thread of execution. This means, that this thread will not be awoken by kernel until timeout completes. What it means, is that none of our fibers will be able to use that thread. This is bad, because usually thread pools are made to fit in-line with number of machine CPU cores. In practice, Thread.Sleep may keep one of our CPU cores idle, wasting machine power in the process.

+

For this reason we usually want to build a suspendable fibers, that will respect our thread pool. This however cannot be done without cooperation with scheduler itself. Therefore, we need to extend API of our scheduler:

+
1
2
3
type IScheduler =
abstract Schedule: (unit -> unit) -> unit
abstract Delay: TimeSpan * (unit -> unit) -> unit
+ +

And our simple implementation of it as well:

+
1
2
3
4
5
6
7
8
9
10
11
let shared = 
{ new IScheduler with
member __.Schedule fn = ....
member this.Delay (timeout: TimeSpan, fn) =
let mutable t = Unchecked.defaultof<Timer>
let callback = fun _ ->
t.Dispose()
fn()
()
t <- new Timer(callback, null, int timeout.TotalMilliseconds, Timeout.Infinite)
}
+ +

We’ll use a .NET timers here to implement our delays. With these in our hands, ourFiber.delay operation is trivial to implement:

+
1
2
3
4
5
6
7
let delay (timeout): Fiber<unit> =
Fiber <| fun (s, c) next ->
if c.Cancelled then next None
else s.Delay(timeout, fun () ->
if c.Cancelled
then next None
else next (Some (Ok ())))
+ +

Composing parallel fibers

We’re slowly getting to the end. What I left for this blog post was to implement two basic operators, that are prevalent in most coroutine libraries:

+
    +
  • Fiber.parallel which will schedule multiple fibers to run in parallel and returns a fiber which aggregates their results.
  • +
  • Running two fibers in parallel and returning the result of whichever completes first, while cancelling a second one. We already discussed this approach before. Here I’ll call it Fiber.race.
  • +
+

Aggregating parallel results

We’ll start from building a parallel operator, which will change our array of fibers into fiber with an array of results. But let’s define the semantics of that operation first:

+
    +
  • Our result fiber completes only when all of the aggregated fibers completed with successful result.
  • +
  • If any of the fibers fails, the resulting fiber also fails.
  • +
  • If any of the fibers fails or get cancelled, all pending ones are also cancelled.
  • +
+

The core skeleton of that operation could look like following:

+
1
2
3
4
5
6
7
8
9
10
let parallel (fibers: Fiber<'a>[]): Fiber<'a[]> =
Fiber <| fun (s, c) next ->
if c.Cancelled then next None
else
let child = c.AddChild()
let successes = Array.zeroCreate remaining
let mutable remaining = Array.length fibs
fibers |> Array.iteri (fun idx (Fiber call) ->
s.Schedule (fun () -> (* to be defined *))
)
+ +

Here, we create a dedicated cancellation token, an array of results and a countdown counter - we’re going to decrement it every time one of our fibers completes to know when we’re ready to return a complete result. I’ve left a placeholder for a lambda body that we actually want to schedule. We’re going to fill it right away:

+
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
// defined above: s.Schedule <| fun () ->
call (s, child) (fun result ->
match result with
| Some (Ok success) ->
// fill the result array
successes.[idx] <- success
if c.Cancelled && Interlocked.Exchange(&remaining, -1) > 0 then
next None
elif Interlocked.Decrement(&remaining) = 0 then
// if all results have been returned, call the `next` callback
if c.Cancelled then next None
else next (Some (Ok successes))
| Some (Error fail) ->
if Interlocked.Exchange(&remaining, -1) > 0 then
child.Cancel() // we failed, cancel other fibers
if c.Cancelled then next None
else next (Some (Error fail))
| None ->
if Interlocked.Exchange(&remaining, -1) > 0 then next None)
+ +

As you probably noticed, we’re using Interlocked class again - that’s because now we have multiple fibers running in parallel, therefore our access to shared mutable values is not thread safe. This includes remaining counter decrement operation. This however doesn’t apply to successes.[i] <- success - since every fiber knows and touches only its own index within result array, there’s no worry that any other will try to push its result in the same place.

+

What you also can see, we’re using a -1 here as a magic value - we’ll use it on the counter as a flag to determine if any of the fibers failed/was cancelled - and if so, which one of them will call the next callback.

+

Racing to completion

With first operator (Fiber.parallel) ready, now it’s the time to implement Fiber.race. Since I’ve discussed it behavior multiple times in this post already, let’s dive straight into the code:

+
1
2
3
4
5
6
7
8
9
10
let race (Fiber left) (Fiber right): Fiber<Choice<'a, 'b>> =
Fiber <| fun (s, c) next ->
if c.Cancelled then next None
else
let mutable flag = 0
let cancelChild = c.AddChild()
let run fiber choice =
(* to be described *)
run left Choice1Of2
run right Choice2Of2
+ +

So again, we want to have shared mutable flag, which we’ll use to determine, which of the fibers finished as a first one to be able to call fiber’s callback safely and cancel the other. You may see, that our returned fiber uses Choice<,> type - this means, that our left and right fibers can have results of different types. We’ll use that soon, but first we need to complete our run function body:

+
1
2
3
4
5
6
7
8
9
10
let run fiber choice =
s.Schedule (fun () ->
fiber (s, cancelChild) (fun result ->
if Interlocked.Exchange(&flag, 1) = 0 then
cancelChild.Cancel()
if c.Cancelled then next None
else match result with
| None -> next None
| Some(Ok v) -> next (Some(Ok(choice v)))
| Some(Error e) -> next (Some(Error e))))
+ +

What we do here is simply trying to race to “reserve” out flag variable - the winner gets his result mapped to corresponding choice, while looser gets cancelled.

+

What’s interesting, we can now combine our race and delay functions to easily implement timeout mechanism:

+
1
2
3
4
5
6
7
8
9
10
11
let timeout (t: TimeSpan) fiber =
Fiber <| fun (s, c) next ->
let (Fiber call) = race (delay t) fiber
call (s, c) (fun result ->
if c.Cancelled then next None
else match result with
| None -> next None
| Some(Ok (Choice1Of2 _)) -> next None // timeout won
| Some(Ok (Choice2Of2 v)) -> next (Some(Ok v))
| Some(Error e) -> next (Some(Error e))
)
+ +

The one last thing left for us, is to be able to run out fibers on the main thread - otherwise we’d start our program, schedule fibers to run in the background and then close the program without waiting for the results.

+
1
2
3
4
5
6
7
8
9
let blocking (s: IScheduler) (cancel: Cancel) (Fiber fn) =
use waiter = new ManualResetEventSlim(false)
let mutable res = None
s.Schedule(fun () -> fn (s, cancel) (fun result ->
if not cancel.Cancelled then
Interlocked.Exchange(&res, Some result) |> ignore
waiter.Set()))
waiter.Wait()
res.Value
+ +

It’s simple - we’ll use standard synchronization primitives provided by .NET runtime, to hold current OS thread until we complete. Sure it’s blocking an OS thread, but we’ll eventually need that if we don’t want our program’s main function to finish before all fibers inside the thread pool complete.

+

Simulating real environment in tests

In theory, we could be done here. But, if you managed to read up to this point, we may want to cover one last scenario. Imagine that we’d want to test our fibers. However running tests using standard thread pool scheduler can lead to funky issues:

+
    +
  • Sometimes you may trigger some race conditions in your code, that only happen in specific situations (like high CPU contention) and are almost impossible to reproduce during debug sessions.
  • +
  • Other times you may have some lengthy delays/timeouts in your code, like waiting for seconds or even minutes before continuing. Guess what: now your test will wait for just as long.
  • +
+

These are not new problems. They are well known in world of concurrent and distributed systems. What we need, is a simulation of execution environment. If you want to listen more about that concept, I could recommend you this presentaton. To run our test predictably, we’ll create a dedicated test scheduler, which will run our code in deterministic fashion (on a single core) and in a way that’s detached from other invariants eg. actual physical clock and random number generator.

+

The idea here is simple - our scheduler will operate on notion of virtual timeline. When we’ll try to schedule a new function - to trigger either immediately or after some timeout - we’ll store it inside an ordered collection, a timeline. Some of the technical decisions we also made for purposes of this implementation:

+
    +
  • Whenever a fiber is going to schedule multiple parallel executions “at the same time”, we’ll put them all into a single bucket on a timeline. Later on I’ll cover, why this is useful.
  • +
  • We’ll assume, that single operation execution is instantaneous. It means, it doesn’t advance our scheduler’s clock. We do it only for delayed executions.
  • +
+

After describing the concept behind the algorithm, the actual implementation really shouldn’t be that surprising:

+
1
2
3
4
5
6
7
8
9
10
11
12
13
type TestScheduler(now: DateTime) =
let mutable running = false
let mutable currentTime = now.Ticks
let mutable timeline = Map.empty
let schedule delay fn = (* to be defined *)
let rec run () = (* to be defined *)
interface IScheduler with
member this.Schedule fn =
schedule 0L fn
if not running then
running <- true
run ()
member this.Delay (timeout: TimeSpan, fn) = schedule timeout.Ticks fn
+ +

We’re using a running flag here to not try to invoke run multiple times in nested manner - this would cause non-tailable recursion and potential stack overflow in more expensive tests.

+

The schedule function is pretty simple - calculate expected execution time for a function, then add that function to be executed at that point in time.

+
1
2
3
4
5
6
let schedule delay fn = 
let at = currentTime + delay
timeline <-
match Map.tryFind at timeline with
| None -> Map.add at [fn] timeline
| Some fns -> Map.add at (fn::fns) timeline
+ +

Given all of the code we already survived in this blog post, run loop should be pretty simple:

+
1
2
3
4
5
6
7
8
9
let rec run () =
match Seq.tryHead timeline with
| None -> running <- false
| Some (KeyValue(time, bucket)) ->
timeline <- Map.remove time timeline
currentTime <- time
for fn in List.rev bucket do
fn ()
run ()
+ +

We’ll try to pick the first entry from the timeline - since here we use F# map, which is sorted in ascending order, we know that first entry is the one with the shortest execution timeout. We update our “current” time to match the expected one we calculated earlier, and finally we execute all functions scheduled at that time and repeat the loop all over until we eventually run out of scheduled actions.

+

Now here’s the trick - we use List.rev to execute functions in the same order in which they were scheduled, because we want our tests to be deterministic and our bugs to be reproducible. However this is not the only strategy - since we know that functions in the same bucket could as well be executing in parallel, we could shuffle them around in different permutations for early discovery of some data races! I’ll won’t dive into it, but leave that idea as food for thoughts for you.

+

One last note about the test scheduler is that isolating it from the actual physical clock means, we cannot trust our time functions (like DateTime.UtcNow) any longer. This shouldn’t really be an issue though - because relying on physical time would potentially make our tests indeterministic, we didn’t want to use it anyway, right?

+

However, we need to be able to obtain current time from the scheduler, so we need to extend its API:

+
1
2
3
4
5
6
7
8
9
10
type IScheduler =
abstract UtcNow: unit -> unit
// ... other methods

type TestScheduler() =
let mutable currentTime = DateTime.UtcNow.Ticks
// ... rest of the implementation
interface IScheduler with
member __.UtcNow() = DateTime(currentTime)
// ... other methods
+ +

And that’s all. As always, if you got confused or have a problems along the way, you can get the entire code here. I wanted to thank to Anthony Lloyd for his initial work on porting Scala ZIO library to F#, which brought me an inspiration to write this piece.

+ +
+ + + + + +
+ + + + + + + +
+ + +
+
+
+ + + +
+ + +
+
+
+
+
+
+ +
+
+
+
+
+
+
+
+
+ + + + + + + + + diff --git a/2024/10/18/Dealing-with-complex-dependency-injection-in-FSharp/index.html b/2024/10/18/Dealing-with-complex-dependency-injection-in-FSharp/index.html new file mode 100644 index 00000000..8195ce32 --- /dev/null +++ b/2024/10/18/Dealing-with-complex-dependency-injection-in-FSharp/index.html @@ -0,0 +1,370 @@ + + + + + + + + + + + + + + + + + + + + + +Dealing with complex dependency injection in FSharp | 暮秋小屋 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+
+ +
+ +
+
+
+ + + +
+
+
+ + +
+
+
+ + +
+ +
+ +
+ +
+
+
+

Today, we’re going to cover different ways of encapsulating capabilities and supplying them between functions using functional programming techniques which can be realized in F#.

+

Managing code dependencies in object oriented languages in 2020 is pretty much one sided problem: dependency injection has won, people use dedicated frameworks to handle that for them, which 99.9% of the time operate using runtime reflection. Of course now you need to learn them as well, potentially misconfigure them and fail at runtime or maybe even discover that not every problem is a stateless web service, but it still better (?) than what we had in the past, and what more can we possibly do anyway?

+

On the other side in functional space, there’s no one opinionated solution or approach - various things have been proposed, usually depending on features that languages and compiler have to offer. And since pretty much all functional languages offer this thing known as partial application, for many years it was the most common answer for the problem of managing dependencies.

+

In short we’re talking about dependency injection by function parameter, like:

+
1
2
3
4
5
6
7
// foo requires 2 dependencies to serve the incoming request
let foo bar baz request = ???
// we're providing dependencies by partial application
let wired = foo dependency1 dependency2
// now wired can serve request directly without calling
// dependencies every time
let response = wired request
+ +

It’s very simple, doesn’t require reflection or dedicated library. However there are several pain points coming with this approach - visible especially as our code base grows and become more complex. However latest approaches popularized by libraries like Scala ZIO or Haskell’s Polysemy challenge this approach.

+

Partial application

There are some design decisions, when partial application doesn’t always give a clear answer. Example:

+
+

Imagine using a set of methods, that are closely related and - in object oriented world - encapsulated within  a single object, like database query/execute or different logging methods (debug/info/warning/error).

+
+

Now, given that our function needs to use potentially more than one of these, how should we pass our arguments?:

+
    +
  1. Functional purist path - pass every dependency as a separate function parameter: let doSmth logError logInfo = ??. While it allows us to precisely describe what this function uses, it would of course lead to explosion of function parameters. Additionally every time you need new function in your existing code, you need to partially apply it at all call sites.
  2. +
  3. Describe operations using ADT (algebraic data types) and inject a function that will work as an interpreter for them: let doSmth (log: LogEvent -> unit) = ??. While it’s easy to mock (you don’t need to implement everything, only pattern match on cases that matter for a particular test) and reduces params affinity, it also comes with a lot of indirection, that may lead to harder to grasp, especially during debugging. Sometimes a performance penalty is also to be expected.
  4. +
  5. Fallback to objects/interfaces and pass them as methods: let doSmth (logger: ILogger) = ??. While interfaces may simplify dependency tree, it’s not always obvious when to use it. Mocking story is also more painful + interfaces are not inferred by F# compiler.
  6. +
+

These are quite common options I’ve seen in the wild - each having their own advantages and disadvantages. Which one to use? Good question, as in practice with codebases that are old enough, you usually see 2 or even all 3 of them mixed together. This can lead to some confusion and obscurity over time.

+

What’s worse, none of these cases really solves problem of dependency management - all they do is just try to reduce it to a manageable scope. Eventually you’ll end up manually wiring - by partial application - dozens of functions and managing all of the dependencies between them by hand.

+

Example

In order to get better understanding, we’ll use a fairly simple example - changing user password:

+
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
let fetchUser (db: IDbConnection) userId = 
db.QueryFirstAsync(Sql.FetchUser, {| userId = userid| })

let updateUser (db: IDbConnection) user = db.ExecuteAsync(Sql.UpdateUser, user)

let changePass (logger: ILogger) fetch update = fun req -> task {
let! user = fetch req.UserId
if user.Hash = bcrypt user.Salt req.OldPass then
let salt = generateSalt ()
let user' = { user with Salt = salt; Hash = bcrypt salt req.NewPass }
do! update user'
logger.LogInformation "Password change: user %i" user.Id
return Ok ()
else
logger.LogError "Password change unauthorized: user %i" user.Id
return Error "Old password is invalid"
}

+ +

Here we have a fairly short snippet with some dependencies? But how many in practice?:

+
    +
  • Number of parameters suggest 3, but depending on our choices it could be 4 (if we decide to pass log error and info separately) or 2 (if we conflate fetch/update into dedicated interface).
  • +
  • It’s not hard to imagine that in the future our bcrypt hashing function may turn out to be configurable - maybe even per each user. That may need a configurable parameter.
  • +
  • Maybe aside of the logger we may be needing a separate telemetry mechanism to count number of incoming request or password validation failures? That means another parameter.
  • +
  • Salt generation is pseudo-random process - it we want our function to be deterministic, we should probably parametrize it over explicitly passed Random as well.
  • +
+

As you see, what seemed to be simple task at the beginning can quickly blow up out of proportion. As the number of arguments grows, the more nasty our wiring code eventually becomes. Quite common pattern is to hide all of that nastiness under the carpet a.k.a. composition root. However this doesn’t have to be the case.

+

Below we’ll cover another approach for dealing with dependencies - inspired by Scala ZIO library - using incremental steps, from first principles to monadic bindings.

+

Managing dependencies beyond partial application

Let’s start from how our code from above will eventually look like at the end of this step:

+
1
2
3
4
5
6
7
8
9
10
11
let changePass env = fun req -> task {
let! user = Db.fetchUser env req.UserId
if user.Hash = bcrypt user.Salt req.OldPass then
let salt = Random.bytes env 32
do! Db.updateUser env { user with Salt = salt; Hash = bcrypt salt req.NewPass }
Log.info env "Changed password for user %i" user.Id
return Ok ()
else
Log.error env "Password change unauthorized: user %i" user.Id
return Error "Old password is invalid"
}
+ +

As you may notice, all of our partially applied parameters disappeared, replaced by some single cryptic env parameter. We’ll get there soon. We also packed similar capabilities into corresponding modules (Db/Log/Random). Lets start from defining them:

+
1
2
3
4
5
6
7
8
9
10
[<Interface>]
type ILogger =
abstract Debug: string -> unit
abstract Error: string -> unit

[<Interface>] type ILog = abstract Logger: ILogger

module Log =
let debug (env: #ILog) fmt = Printf.kprintf env.Logger.Debug fmt
let error (env: #ILog) fmt = Printf.kprintf env.Logger.Error fmt
+ +

Now we can say something more about env. The secret is in #ILog signature, which means that our environment can be any generic type implementing ILog interface. As soon you’ll see, this approach is highly composable, but before that we’ll need another module:

+
1
2
3
4
5
6
7
8
9
10
11
[<Interface>]
type IDatabase =
abstract Query: string * 'i -> Task<'o>
abstract Execute: string * 'i -> Task

[<Interface>] type IDb = abstract Database: IDatabase

module Db =
let fetchUser (env: #IDb) userId =
env.Database.Query(Sql.FetchUser, {| userId = userId |})
let updateUser (env: #IDb) user = env.Database.Execute(Sql.UpdateUser, user)
+ +

Now what will happen if we use functions from both Log and Db modules? As it turns out, F# compiler can properly infer generic constraints over these interfaces. The result env type constraint is inferred to be an union - just like set union, which also means that it handles duplicates for us - of all constraints of other functions using env in its scope:

+
1
2
3
let foo env = // env :> IDb and env :> ILog
let user = Db.fetchUser env 123 // env :> IDb
Log.debug env "User: %A" user // env :> ILog
+ +

Now why did we use two separate interfaces (ILog/ILogger) instead of making environment implement ILogger directly? This is more practical approach that will let us isolate capabilities of particular modules rather than putting them flat into our environment. Example:

+
1
2
3
4
5
6
7
8
9
module Log =
let live : ILogger = ?? // create logger interface

[<Struct>]
type AppEnv =
interface ILog with member _.Logger = Log.live
interface IDb with member _.Database = Db.live connectionString

foo (AppEnv())
+ +

We cannot eagerly provide a specific implementation of ILog/IDb, because they’re yet to be defined as part of by our environment type (which may need to implement many interfaces). To maintain module encapsulation Log module shouldn’t be aware of existence or implementation of IDb interface and vice versa for Db module. What we can do however is to provide live implementation of ILogger, which encapsulates capabilities required by the Log module. This way we don’t need to know details of ILogger when defining our environment type.

+

Strong sides of this approach?:

+
    +
  • We only need to provide a single environment object instead of (potentially) unbounded list of parameters. Since it’s always one, it’s easier to generalize and compose other functions over it.
  • +
  • Unlike in reflection-based dependency injection frameworks - everything is still safe and checked by the compiler. If our environment type will not implement an interface required somewhere in the call chain, our code will simply not compile.
  • +
  • It’s still fairly easy to unit test - each function defines only the generic type constraints that it uses in its own call tree, NOT all of the constraints required by the application.
  • +
  • New dependencies are added implicitly - if your code uses module that requires additional capability, it will be automatically inferred by the compiler and bubble up to our environment type definition. No need to add new function parameters or to pass new argument. Also - unlike the object oriented IoC containers - there’s no need to add new dependency as a field or constructor argument.
  • +
  • It gives some opinionated approach on what should be a dependency - less thinking of “should that be a function or interface?” or “if these two functions correspond to the same capability, should they be passed separately?”, which arguably may be a good thing.
  • +
  • It doesn’t impose specific restrictions on libraries and frameworks.
  • +
+

Now we could as well stop here - IMHO this approach is already good and useful for most cases. We can also try to push it further. As you’ve seen, our code now requires quite a lot of env passing around. Could we do something about this? It turns out that yes, we could.

+

Reader monad

Before we continue: what we’re going to cover now is less useful in terms of current state of F# ecosystem for the reasons I’ll mention later.

+

The pattern we’ll use here is known as a Reader Monad. While it’s useful in certain situations, it’s not widely used - IMO it’s fault lies in the name itself, which somehow managed to sound both borderline meaningless and scary in ears of many developers.

+

The rest of this blog post will be introduction to this style in F#, however focused solely around problem of dependency management - we’ll ignore other aspects of monads.

+

We’ll going to reuse our environment type from above, but now encode it directly into another type we’ll call Effect. Since I’ve mentioned that our pattern has M-word in it, you can safely assume that our handler’s logic will be defined as a lazy sequence of steps to be executed (sounds almost like async/await). In F# we’ll sugar them by using custom computation expression (I’m going to call it effect { ... }) returning our effect type, which we’ll define as:

+
1
[<Struct>] type Effect<'env, 'out> = Effect of ('env -> 'out)
+ +

Where:

+
    +
  • env is our environment type we already talked about above.
  • +
  • out defines a returned value type of our effect.
  • +
+

Eventually, with this type in hand our simple request handler will be looking like that:

+
1
2
3
4
5
6
7
8
9
10
11
let changePass req = effect {
let! user = Db.fetchUser req.UserId
if user.Hash = bcrypt user.Salt req.OldPass then
let! salt = Random.bytes 32
do! Db.updateUser { user with Salt = salt; Hash = bcrypt salt req.NewPass }
do! Log.info "Changed password for user %i" user.Id
return Ok ()
else
do! Log.error "Password change unauthorized: user %i" user.Id
return Error "Old password is invalid"
}
+ +

As you see, there’s no more env parameter being passed around. It’s now an implicit part of our effect expression. However at the moment we didn’t provide enough infrastructure in our code to make that thing work. What we’re going to need is a set of operators, we can use to make our computation expression happen.

+

First we’re going to need some Effect<'env,'out> constructors:

+
1
2
3
4
5
module Effect =
/// Create value with no dependency requirements.
let inline value (x: 'out): Effect<'env,'out> = Effect (fun _ -> x)
/// Create value which uses depenendency.
let inline apply (fn: 'env -> 'out): Effect<'env,'out> = Effect fn
+ +

We also need some way to run our effect to be able to make it… well effectful:

+
1
2
3
4
5
module Effect =

(* ...other functions... *)

let run (env: 'env) (Effect fn): 'out = fn env
+ +

And since we already mentioned Effect is monad, we also gonna need a bind function as well if we want to compose our effects together:

+
1
2
3
4
5
6
7
8
9
module Effect =

(* ...other functions... *)

let inline bind (fn: 'a -> Effect<'env,'b>) effect =
Effect (fun env ->
let x = run env effect // compute result of the first effect
run env (fn x) // run second effect, based on result of first one
)
+ +

This is pretty much it. We’re just going to add compose all of these into builder type to make it usable as F# computation expression:

+
1
2
3
4
5
6
7
8
[<Struct>]
type EffectBuilder =
member inline __.Return value = Effect.value value
member inline __.Zero () = Effect.value (Unchecked.defaultof<_>)
member inline __.ReturnFrom (effect: Effect<'env, 'out>) = effect
member inline __.Bind(effect, fn) = Effect.bind fn effect

let effect = EffectBuilder()
+ +

Of course this, we still need to adapt the modules we prepared earlier to now operate on effects rater than plain functions. We can make this easier by using our Effect.apply function, like:

+
1
2
3
4
module Log =    
let debug fmt =
let ap s = Effect.apply (fun (x: #ILog) -> x.Logger.Debug s)
Printf.kprintf ap fmt
+ +

So - as you may have noticed in final form of our effect-based changePass function - in result we almost fully erased all of the dependency-wiring code from our example. There are several downsides of this approach:

+
    +
  • We do a lot of more bindings (see let!/do! expressions), which means more lambda closures, indirection (wait to see call stacks) and more allocations.
  • +
  • Altogether we also erased task { ... } computation expression and with it an out-of-the-box ability to write asynchronous code. This is one of the downsides of using monads - cross-type composition is painful.
  • +
+

Of course we could enrich our Effect type to be able to bind it with Task/Async. That however means, that our pattern grows in complexity and becomes more of a framework rather than something to be easily applied into existing code. Is that bad? Not necessarily, but for sure comes with a bigger commitment, as now you’re not only writing business logic but eventually maintain new effect library. Maybe in future this concept will grow into its own space in favor of the F# ecosystem.

+

Summary

We came from partial application as tool for dependency injection, over more structured approach promoting single environment type with help of powerful F# type inference, up to encapsulating it into a Reader Monad. That’s a long way. I hope you’ll give it a try and it will help you determine the approach that works for you.

+ +
+ + + + + +
+ + + + + + + +
+ + +
+
+
+ + + +
+ + +
+
+
+
+
+
+ +
+
+
+
+
+
+
+
+
+ + + + + + + + + diff --git a/2024/10/18/Generalised-signature/index.html b/2024/10/18/Generalised-signature/index.html new file mode 100644 index 00000000..b5297985 --- /dev/null +++ b/2024/10/18/Generalised-signature/index.html @@ -0,0 +1,383 @@ + + + + + + + + + + + + + + + + + + + + + +Generalised signature | 暮秋小屋 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+
+ +
+ +
+
+
+ + + +
+
+
+ + +
+
+
+ + +
+ +
+ +
+ +
+
+
+

#ocaml #fp

+
+

This post presents a technique for defining more reusable OCaml signatures, helping to maintain consistent APIs with minimal boilerplate. We’ll work through a few examples, which you can check out on GitHub.

+
+

Indexable containers

Consider the following definition of an iter function for some container type t:

+
1
2
3
4
let iter f t =
for i = 0 to length t - 1 do
f (get t i)
done
+ +

iter requires only that t comes with functions get and length. Many useful operations can be derived in terms of such indexing functions. To take advantage of this, let’s move iter into a functor and provide some other useful operations too:

+
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
module type Indexable1 = sig
type 'a t

val get : 'a t -> int -> 'a
val length : _ t -> int
end

module Foldable_of_indexable1 (I : Indexable1) : sig
open I

val iter : ('a -> unit) -> 'a t -> unit
val iteri : (int -> 'a -> unit) -> 'a t -> unit
val fold_left : ('acc -> 'a -> 'acc) -> 'acc -> 'a t -> 'acc
val exists : ('a -> bool) -> 'a t -> bool
val for_all : ('a -> bool) -> 'a t -> bool
val is_empty : _ t -> bool
(* ... *)
end
+ +

For many types, including array, the get-based definitions are identical to their hand-optimised equivalents (modulo functor application). We can imagine avoiding a lot of standard-library boilerplate – and potential for API inconsistency – by using many such functors 1. We’d end up defining exactly one iter function that suffices for all Indexable types.

+

All good so far. Now, let’s consider the string type.

+

A string is also an indexable container with length and get functions, albeit one that can only contain char values. It’s natural to expect to be able to re-use Foldable_of_indexable1 in some way: indeed, our definition of iter above is exactly equal to the one in Stdlib.String.iter. Unfortunately, our Indexable1 module type can only describe parametric containers:

+
1
module _ : (Indexable1 with type 'a t := string) = Stdlib.String
+ +
1
2
3
4
5
6
7
Error: Signature mismatch:
...
Values do not match:
val get : t -> int -> char
is not included in
val get : t -> int -> 'a
File "string.mli", line 52, characters 0-57: Actual declaration
+ +

We’re unable to tell the type system something like

+
+

'a t = string    implies    'a = char

+
+

as part of our substitution. This means that many types – including string, bytes, unboxed arrays and unboxed vectors – can’t benefit from our Foldable_of_iterable1 definitions, even though their own definitions will be identical!

+

When we wrapped our code in the Foldable_of_indexable1 functor, we needed to give it specific input and output module types, and the ones we picked artificially limited its usefulness. This is a hazard of functorising highly-generic code. As ever, we could solve the problem with copy-paste: a new Indexable0 module type for non-parametric containers, and a new functor Foldable_of_indexable0 with exactly the same implementations as our previous one.

+
1
2
3
4
5
6
7
8
9
10
11
12
(* Non-parametric indexable types *)
module type Indexable0 = sig
type t
type elt

val get : t -> int -> elt
val length : t -> int
end

module Foldable_of_indexable0 (I : Indexable0) : sig
(* All with the same implementation as before... *)
end
+ +

This definition suffers from the dual problem when we try to apply it to parameterised containers like 'a array:

+
1
module _ : (Indexable0 with type t := 'a array) = Stdlib.Array
+ +
1
Error: The type variable 'a is unbound in this type declaration.
+ +

This time, we wanted to be able to say something like

+
+

elt = 'a    implies    t = 'a array    (where 'a is universally quantified),

+
+

which is even more nonsensical than our previous attempt. Neither Indexable0 nor Indexable1 can be expressed in terms of the other. We need something more general.

+

Something more general

Interestingly, it’s possible to generalise Indexable0 and Indexable1 with another layer of indirection by making elt a type operator:

+
1
2
3
4
5
6
7
module type IndexableN = sig
type 'a t
type 'a elt

val get : 'a t -> int -> 'a elt
val length : _ t -> int
end
+ +

elt carries the type equalities needed for the Indexable1 case, without forbidding the non-parametric implementation needed for the Indexable0 case. Arrays can set 'a elt := 'a, and strings can set 'a elt := char. Indeed, we can do this in the general case:

+
1
2
3
4
5
6
7
8
9
(** [Indexable0] is a special-case of [IndexableN] *)
module Indexable0_to_N = functor
(T : Indexable0) ->
(T : IndexableN with type 'a t := T.t and type 'a elt := elt)

(** [Indexable1] is a special-case of [IndexableN] *)
module Indexable1_to_N = functor
(T : Indexable1) ->
(T : IndexableN with type 'a t := 'a T.t and type 'a elt := 'a)
+ +

Now we can define a single Foldable_of_indexableN functor (with exactly the same implementations as before), and it will work for polymorphic and monomorphic containers. Neat!

+

A lattice showing Indexable0 and Indexable1 being generalised by IndexableN.

+

In the general case, when you notice that different signatures are sharing common functions, it’s often possible to unify them under a common interface with the following two steps:

+
    +
  1. generalise. Convert pure type variables into type operators (as in 'a → 'a elt), to support use-cases like instantiating those variables to fixed types. Add type parameters to existing types to carry type equalities between them (as in 'a t / 'a elt), to support use-cases where these types depend on each other.

    +
  2. +
  3. specialise. Use destructive substitution (:=) to eliminate those types and type parameters when they’re not needed. We’re taking advantage of the more powerful destructive substitution offered by OCaml 4.06, which allows us to freely undo our generalisation step.

    +
  4. +
+

The truly magical part of this trick is that – with better support for destructive type substitutions recently added to Odoc – it can be made completely invisible2 in documentation!

+
1
2
3
4
5
6
7
8
9
10
11
module type Indexable1 = sig
type _ t

val get : 'a t -> int -> 'a
val length : _ t -> int
end

(** This module gets identical documentation to the one above! *)
module type Indexable1' = sig
include IndexableN with type 'a elt := 'a (** @inline *)
end
+ +

What’s the cost?

One unavoidable limitation is in what sort of operations we can put in the Foldable_of_indexable functor. Suppose our initial attempt at generalising containers included a sum function:

+
1
let sum : int t -> t = fold_left ( + ) 0
+ +

sum requires a container that can hold int values, which is clearly not possible for strings as the type system will happily tell us:

+
1
2
3
4
5
   |   let sum = fold_left ( + ) 0
^^^^^
Error: This expression has type int -> int -> int
but an expression was expected of type int -> 'a elt -> int
Type int is not compatible with type 'a elt
+ +

To state the obvious, we can’t rely on parametricity in our container functions if we want them to work on non-parametric containers. The natural solution here would be to define such parametric-only functions in a separate functor.

+

Other examples

Indexable containers aren’t the only example of generalised signatures in the real world. Indeed, many other data-structures and design patterns have APIs that can be unified in this way. Consider the case of hashtables, which have a huge space of possible implementations:

+
    +
  • key types can be left polymorphic by using a magic hash function like caml_hash (as in Stdlib.Hashtbl), or fixed by a user-specified hash function (as in Stdlib.Hashtbl.Make).

    +
  • +
  • value types can be left polymorphic, fixed by the user (as in persistent hashtables like Index), or even determined by the keys used to index them (as in universal maps like Hmap).

    +
  • +
+

Initially, it looks like these different hashtables will each require their own hand-written signature (and this is what the standard library does with its hashtables). However, with enough type parameters, these different implementations can all be unified under a single Hashtbl_generalised module type:

+
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
module type Hashtbl_generalised = sig
(** We have three types ([t], [key] and [value]) and three type variables:

- ['k]/['v] allow the hashtable to determine key/value types;
- ['a] is carried from keys to corresponding values, allowing the key to
determine the types of values. *)

type ('k, 'v) t
type ('k, 'a) key
type ('v, 'a) value

val create : int -> (_, _) t
val replace : ('k, 'v) t -> ('k, 'a) key -> ('v, 'a) value -> unit
val remove : ('k, _) t -> ('k, _) key -> unit
val find_opt : ('k, 'v) t -> ('k, 'a) key -> ('v, 'a) value option
(* ... *)
end
+ +

We can then implement our different hashtable signatures as specialisations:

+

A lattice showing four different `Hashtbl` module types being generalised by `Hashtbl_generalised`.

+

For instance, for the regular polymorphic hashtable:

+
1
2
3
4
5
module type Poly_hash = sig
include Hashtbl_generalised
with type ('k, _) key := 'k
and type ('v, _) value := 'v (** @inline **)
end
+ +

The other specialisations are very similar (see here for the specifics).

+

What is it that makes Hashtable_generalised a good parent interface for these four flavours of hashtable? To get some insight, we can notice that each of the type parameters ('k, 'v, and 'a) connects its own pair of types:

+

hashtbl_generalised

+

+

Framed this way, the type parameter 'k exists solely to carry type information between hashtables and their keys (using a type equality at call sites). Similarly, 'v bridges between hashtables and values, and 'a between keys and values. From here, each of our hashtable variants uses destructive subsitution (:=) to prune away unnecessary bridges and express some sort of dependency relation between the types:

+ + + + + + + + + + + + + + + +
poly_hash



Keys and value types constrain t at call-sites.
mono_hash



The value type constrains t, but keyis fixed by a functor.
persistent



Both key and value types are fixed by a functor.
universal



Each key’s type constrains the corresponding value’s type.
+

In this case, it’s not feasible for all these data structures to share the same implementation, but it’s still valuable for them to implement a common core API: it ensures consistency of the user-facing functions, allows sharing of documentation, and may even allow these implementations to share a common test suite.

+

Conclusion

The full code for our Indexable and Hashtbl examples, including explicit definitions of each of the module types, can be found in the generalised-signatures repository. This repository also contains and a third demonstration of this technique being used to express monad-like signatures. The auto-generated documentation for these examples can be viewed online.x

+

Thanks for making it to the end; I hope you picked up something useful. If you think it would help others in your network, I’d appreciate it if you shared it with them.

+
+

Appendix A: Haskell suffers too

The typeclasses in Haskell’s base have the same “polymorphic-instances-only” property as our Indexable1 signature (unsurprising, since it doesn’t provide any unboxed container types).

+
1
2
3
class Indexable1 f where        -- Polymorphic instances only
get :: f a -> Int -> a
length :: f a -> Int
+ +

A similar trick can be performed there to generalise the typeclass instances for monomorphic containers like Text:

+
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
{-# LANGUAGE TypeFamilies #-}

type family Elt container -- Relate containers to their element type
type instance Elt [a] = a
type instance Elt Text = Char

class IndexableN c where
get :: c -> Int -> Elt c
length :: c -> Int

instance IndexableN [a] where -- Polymorphic instance
get = (!!)
length = Prelude.length

instance IndexableN Text where -- Monomorphic instance
get = Text.index
length = Text.length
+ +

As in the OCaml version, we use an Elt type operator to carry the equality needed for the monomorphic case. This time we used type families to specify the relations explicitly, but we could have used multi-parameter type classes for something more akin to the OCaml functor implementation. See the mono-traversable package for more of this sort of trickery in Haskell.

+
+
    +
  1. This is the approach taken by Jane Street’s base, and is very similar to the Haskell notion of building standard libraries from type-class instances.↩
  2. +
  3. This example uses the (** @inline *) tag to ensure that Odoc doesn’t leak that Indexable1' is implemented in terms of IndexableN.↩
  4. +
+ +
+ + + + + +
+ + + + + + + +
+ + +
+
+
+ + + +
+ + +
+
+
+
+
+
+ +
+
+
+
+
+
+
+
+
+ + + + + + + + + -- cgit v1.2.3