R Packages for Knowledge Space Theory

Version 3.0-2

About this document

  • This work is published under the CC-BY-NC-SA license

  • To print this document, you may also use its PDF, DOCX, or ODT forms.

  • This is a living document; please regard the version number, especially in citations.

1 Introduction

Knowledge space theory (KST; Doignon and Falmagne (1985); Doignon and Falmagne (1999)) offers a set-theoretical framework to structure domains of knowledge through prerequisite relationships. These structures allow an efficient adaptive assessment and training of learners’ knowledge.

Since the development of the first R package on KST around 2008 (kst; Stahl (2008); Stahl et al. (2026)), a multitude of R packages has been developed which are described in this paper and most of which are available from CRAN. Figure 1 gives an overview of the currently available R packages and their dependencies and connections.

A diagram showing the dependencies between the various R packages

Figure 1.1: Network of R packages for knowledge space theory

The solid blue arrows depict package imports, the dashed blue arrows package suggestions. The green lines show that packages share S3 object classes.

Chapter 2 gives a short introduction to knowledge space theory (KST). In Chapter 3, packages providing rather general functionalities for basic KST are presented, in Chapter 4, packages for KST extensions follow, namely for competence-based knowledge space theory (CbKST) and for probabilistic knowledge structures. In Chapter 5, packages for more specific purposes are mentioned. Throughout the paper, we will use the phsg and xpl data sets provided by the kstMatrix package.

2 Knowledge Space Theory

Knowledge space theory (KST; Doignon and Falmagne (1985); see also Heller and Stefanutti (2024)) was originally developed as a behavioristic framework supporting efficient adaptive assessment of knowledge where efficiency aims at a reduction of the number of questions to be asked. Opposed to other test-theoretical approaches, KST delivers a non-quantitative result, i.e. the subset of test problems the testee is able to solve instead of a less expressive test score.

Please note that terminology and notations have changed over time, especially in the early years of KST development. In this paper, the line of Doignon and Falmagne (1999) is followed.

In KST, a domain of knowledge is specified as a set \(Q\) of dichotomous test items. The knowledge state of a person is given by the subset \(K\subseteq Q\) of items this person can solve. Usually, not all subsets of \(Q\) can be observed as knowledge states because there exist prerequisite relationships between the items. The set \(\mathcal{K}\) of all possible knowledge states is called knowledge structure. A knowledge structure always contains the emptyset \(\emptyset\) and the full item set \(Q\) as knowledge states.

A special case of a knowledge structure is the knowledge space, a knowledge structure that is closed under union, i.e. for any two knowledge states \(K,K'\in\mathcal{K}\), we have \(K\cup K'\in\mathcal{K}\). If \(\mathcal{K}\) is additionally closed under intersection, it is called quasi-ordinal knowledge space.

A knowledge space can also represented by its basis \(\mathcal{B}\), i.e. the set of all those states which cannot be constructed as union of other, smaller states. The basis is the smallest family of sets from which a knowledge space can be constructed by closure under union. Sometimes, also the term base is used for the basis.

Another basic concept in KST is the surmise relation \(\preceq\). We write \(a\preceq b\) for two items \(a,b\in Q\) if, from a person mastering item \(b\), we can surmise that this person also masters item \(a\). Sometimes, however, this is not so unique because items may be solvable with different solution paths involving different prerequisites. This is covered by the surmise function \(\sigma\), a generalization of the surmise relation. \(\sigma\) is defined by \[\sigma: Q \rightarrow 2^{2^Q\setminus\emptyset}\setminus\emptyset,\] i.e. it assigns to each item \(q\in Q\) a non-empty set of non-empty subsets \(C\subseteq Q\), the clauses of \(q\). If a person masters item \(q\), we can surmise that they also master all items in at least one clause of \(q\). If, for all \(q\in Q\), we have \(|\sigma(q)|=1\), i.e. each item has exactly one clause, the surmise function can be represented as a surmise relation. We have then \(q\preceq q'\) if \(q\in C\in \sigma(q') = \{C\}\).

Both, surmise functions and surmise relations have certain properties: reflexivity, transitivity and (for surmise functions) incompatibility, i.e. any two different clauses for an item are not in a subset relation. From mathematical point of view, a surmise relation is a quasi-order.

There is a one-to-one correspondence between the knowledge spaces over a set \(Q\) and the surmise functions over \(Q\). Similarly, there is a one-to-one correspondence between the quasi-ordinal knowledge spaces and the surmise relations on \(Q\).

2.1 CbKST

While KST itself is rather behavioristic, there was soon a growing interest of KST researchers in the skills and competencies underlying the observable behavior (see, e.g., Doignon (1994); Düntsch and Gediga (1995); Korossy (1997); Albert and Lukas (1999); Heller and Stefanutti (2024)). Out of the various approaches, that of a skill map has prevailed.

A skill map is a triple \((Q,S,\tau)\) where \(Q\) is a non-empty set of problems, \(S\) is a non-empty set of skills, and \(\tau:Q\rightarrow 2^S\setminus\emptyset\) is a mapping assigning to each problem \(q\in Q\) a non-empty set of skills. There are two different interpretations for \(\tau\):

  • Conjunctive model: All skills \(s\in\tau(q)\) are necessary for solving item \(q\in Q\), and
  • Disjunctive model Any skill \(s\in\tau(q)\) is sufficient for solving item \(q\in Q\).

For the conjunctive model, the knowledge state \(K\) of a person in the skill state \(T\subseteq S\) is delineated by \[K=\{q\in Q\mid \tau(q)\subseteq T\}.\] The collection of all knowledge states delineated by some \(T\subseteq S\) is the knowledge structure delineated by the skill map \((Q,S,\tau)\). In the conjunctive model, this knowledge structure is closed under intersection; in the disjunctive model, the delineated knowledge structure is closed under union.

Like surmise relations in classical KST, skill maps do not provide for test items with different solution paths involving different prerequisites. The solution for this are skill multimaps. A skill multimap is a triple \((Q,S,\mu)\) where \(\mu\) is a mapping assigning to each \(q\in Q\) a non-empty collection of non-empty subsets of \(S\). The subsets \(C\in\mu(q)\) are called competencies. For skill multimaps, there is only one interpretation: In order to be able to solve item \(q\), a person must possess all skills \(s\in C\) for at least one competency \(C\in\mu(q)\).

The knowledge state delineated by a subset \(T\subseteq S\) of skills is given by \[K = \{q\in Q\mid C\subseteq T \mbox{ for some }C\in\mu(q)\}.\]

2.2 Probabilistic knowledge spaces

Pure KST is deterministic. However, this is unrealistic when regarding students’ response behavior to test items. Careless errors and lucky guesses occur. In probabilistic KST, this is treated with the Basic Local Independence Model (BLIM). The assumption underlying he BLIM is that this noise depends only on the individual items. Starting with an estimated probability distribution over the knowledge structure (i.e. the probability for each knowledge state to be the learner’s state), the probability to observe a response pattern \(R\subseteq Q\) given a knowledge state \(K\in\mathcal{K}\) is \(P(R) = P(K)\cdot P(R|K)\). If \(\beta_q\) denotes the probability for a careless error answering item \(q\) and \(\eta_q\) denotes the probability for a lucky guess, this leads to \[P(R|K) = \prod_{q\in R\cap K}(1-\beta_q)\cdot\prod_{q\in K\setminus R}\beta_q\cdot\prod_{q\in R\setminus K}\eta_q\cdot\prod_{q\in\bar{R}\cap\bar{K}}(1-\eta_q).\] On the one side, the BLIM can be used for probabilistic knowledge assessment (Falmagne and Doignon (1988a)). On the other side, it can also be used for simulating response patterns. The identifiability of the BLIM is a current research issue.

3 General functionalities

3.1 kst

The kst package (Stahl et al. 2026; Stahl 2008) is a kind of starting point offering basic functionalities for the work with knowledge spaces. The implementation is insofar close to the theory as it uses the sets and relations packages by David Mayer. The kst package uses S3 classes (i) to make sure the corect object types are passed to the functions and (ii) to be able to use generic methods like print() or plot(). Figure 2 shows the hierarchy of object classes in the kst package.

Object classes in the `kst`package

Figure 3.1: Object classes in the kstpackage

All classes are sub classes of set. A core class is kfamset which describes a family of sets. Special cases are kbasis and kstructure, the latter again having a special case in kspace. Somewhat separated is the lpath class describing a learning path.

Working with kst, the first step for us is to convert our example into set notation.

kf <- as.famset(phsg$basis)
kf
#> {{"I1"}, {"I2"}, {"I3"}, {"I2", "I5"}, {"I3", "I6"},
#>  {"I1", "I2", "I3", "I4"}, <<set(7)>>}

By default, sets with more than five elements are not fully printed as can be seen above. This behavior can be changed by passing a limit parameter to the print() or plot() functions.

plot(kf, main="", limit=8)
Plotting a family fo sets with the `kst` package

Figure 3.2: Plotting a family fo sets with the kst package

We can then stepwise derive a knowledge space and, from that, a basis.

ks <- kspace(kstructure(kf))
print(ks, limit=8)
#> {{}, {"I1"}, {"I2"}, {"I3"}, {"I1", "I2"}, {"I1",
#>  "I3"}, {"I2", "I3"}, {"I2", "I5"}, {"I3", "I6"},
#>  {"I1", "I2", "I3"}, {"I1", "I2", "I5"}, {"I1", "I3",
#>  "I6"}, {"I2", "I3", "I5"}, {"I2", "I3", "I6"},
#>  {"I1", "I2", "I3", "I4"}, {"I1", "I2", "I3", "I5"},
#>  {"I1", "I2", "I3", "I6"}, {"I2", "I3", "I5", "I6"},
#>  {"I1", "I2", "I3", "I4", "I5"}, {"I1", "I2", "I3",
#>  "I4", "I6"}, {"I1", "I2", "I3", "I5", "I6"}, {"I1",
#>  "I2", "I3", "I4", "I5", "I6"}, {"I1", "I2", "I3",
#>  "I4", "I5", "I6", "I7"}}
kbase(ks)
#> {{"I1"}, {"I2"}, {"I3"}, {"I2", "I5"}, {"I3", "I6"},
#>  {"I1", "I2", "I3", "I4"}, <<set(7)>>}

The lpath() function determines all learning paths from the empty set to the full item set as a list of the individual paths. The list is rather long, therefore only the first element is printed.

lp <- lpath(ks)
length(lp)
#> [1] 66
print(lp[[1]], limit=8)
#> {{}, {"I1"}, {"I1", "I2"}, {"I1", "I2", "I3"}, {"I1",
#>  "I2", "I3", "I4"}, {"I1", "I2", "I3", "I4", "I5"},
#>  {"I1", "I2", "I3", "I4", "I5", "I6"}, {"I1", "I2",
#>  "I3", "I4", "I5", "I6", "I7"}}

kst offers several further functions including the computation of neighborhoods and fringes and a deterministic knowledge assessment.

3.2 kstMatrix

Work on the kstMatrix package (Hockemeyer et al. 2026; Steiner et al. in press) started originally having a re-implementation of the kst package in mind. The underlying idea was that applying a matrix representation should be faster than the set representation because the former is more native to R. Meanwhile, the package has grown considerably beyond that original purpose into a larger set of functions.

Like in kst, object classes are used intensively as shown in Figure 4.

Object classes in the `kstMatrix` package

Figure 3.3: Object classes in the kstMatrix package

Most classes are derived from the matrix class. Our list starts with those classes that are similar to the ones from the kst package.

  • kmfamset: Family of sets, subclass of matrix
  • kmbasis: Basis of a knowledge space, subclass of kmfamset
  • kmstructure: Knowledge structure, subclass of kmfamset
  • kmspace: Knowledge space, subclass of kmstructure
  • kmqspace: Quasi-ordinal knowledge space, subclass of kmspace
  • kmneighbourhood: Neighborhood of a state in a knowledge structure, subclass of kmfamset
  • kmdata: Data set, subclass of matrix
  • kmattributionrelation: Attribution relation, subclass of matrix
  • kmsurmiserelation: Surmise relation, subclass of kmattributionrelation
  • kmlearningmathmatrix: Learning path as matrix, subclass of kmfamset
  • kmlearningpathmatrices: List of learning paths in matrix form, subclass of list
  • kmlearningpath: Learning path as list of character strings, subclass of list
  • kmlearningpaths: List of learning paths, subclass of list

3.2.1 Basic functions

We start similar as in the description of the kst package above.

phsg$basis
#>      I1 I2 I3 I4 I5 I6 I7
#> [1,]  1  0  0  0  0  0  0
#> [2,]  0  1  0  0  0  0  0
#> [3,]  0  0  1  0  0  0  0
#> [4,]  1  1  1  1  0  0  0
#> [5,]  0  1  0  0  1  0  0
#> [6,]  0  0  1  0  0  1  0
#> [7,]  1  1  1  1  1  1  1
#> attr(,"class")
#> [1] "kmbasis"  "kmfamset" "matrix"   "array"
kmprettyprint(phsg$basis)
#> [1] "{{I1}, {I2}, {I3}, {I2, I5}, {I3, I6}, {I1, I2, I3, I4}, {I1, I2, I3, I4, I5, I6, I7}}"

Prettyprinting can also be done in a simplified way, i.e. omitting braces etc. However, this is recommended only for single character item names.

kmprettyprint(xpl$basis)
#> [1] "{{a}, {b}, {a, c}, {b, c}, {a, b, d}}"
kmprettyprint(xpl$basis, simplified=TRUE)
#> [1] "{a, b, ac, bc, abd}"

Plots with kstMatrix look by default different depending on the class of the plotted object. Various parameters of vertices and edges can be specified in the function call.

plot(phsg$basis)
Plotting a basis with `kstMatrix`

Figure 3.4: Plotting a basis with kstMatrix

3.2.2 Neighborhoods, fringes, and learning paths

The n-neighborhood of a knowledge state is defined as the family of those knowledge states which have a maximal symmetric set difference. By default, the 1-neighborhood is regarded whcih is simply called neighborhood.

ksp <- kmspace(phsg$basis)
state = c(1,1,1,1,0,0,0)
kmneighbourhood(state, ksp, include=TRUE)
#>      I1 I2 I3 I4 I5 I6 I7
#> [1,]  1  1  1  0  0  0  0
#> [2,]  1  1  1  1  1  0  0
#> [3,]  1  1  1  1  0  1  0
#> [4,]  1  1  1  1  0  0  0
#> attr(,"class")
#> [1] "kmneighbourhood" "kmfamset"        "matrix"         
#> [4] "array"
kmnneighbourhood(state, ksp, distance=2)
#>       I1 I2 I3 I4 I5 I6 I7
#>  [1,]  1  1  0  0  0  0  0
#>  [2,]  1  0  1  0  0  0  0
#>  [3,]  0  1  1  0  0  0  0
#>  [4,]  1  1  1  0  0  0  0
#>  [5,]  1  1  1  0  1  0  0
#>  [6,]  1  1  1  1  1  0  0
#>  [7,]  1  1  1  0  0  1  0
#>  [8,]  1  1  1  1  0  1  0
#>  [9,]  1  1  1  1  1  1  0
#> attr(,"class")
#> [1] "kmneighbourhood" "kmfamset"        "matrix"         
#> [4] "array"
plot(kmneighbourhood(state, ksp, include=TRUE), state=state)
Plot of a neighborhood

Figure 3.5: Plot of a neighborhood

Alternatively, a neighborhood an also be marked within the plot of a knowledge structure.

plot(ksp, state=state)
Neighborhood within a knowledge structure

Figure 3.6: Neighborhood within a knowledge structure

The fringe of a knowledge state is the set of items by which the state differs from its 1-neighbors. We distinguish between the inner fringe which is part of the state and the outer fringe containing the items by which the state can be expanded.

kmfringe(state, ksp)
#> I1 I2 I3 I4 I5 I6 I7 
#>  0  0  0  1  1  1  0
kminnerfringe(state, ksp)
#> [1] 0 0 0 1 0 0 0

Learning paths are ways to get from the empty set (of knowing nothing in the field) to the full item set (of knowing everything). In our example space, we have 66 learning paths.

length(kmlearningpaths(ksp))
#> [1] 66
lpm <- kmlearningpathmatrices(ksp)[[1]]
cat(str_wrap(kmprettyprint(lpm), width=60))
#> {} →{I1} →{I1, I2} →{I1, I2, I3} →{I1, I2, I3, I4} →{I1, I2,
#> I3, I4, I5} →{I1, I2, I3, I4, I5, I6} →{I1, I2, I3, I4, I5,
#> I6, I7}
plot(lpm, structure=ksp)
Learning path within a knowledge structure

Figure 3.7: Learning path within a knowledge structure

3.2.3 Deriving and combining knowledge structures

kmunion() and kmintersection() determine the union and intersection, respectively, of knowledge structures. Please note that the union of knowledge spaces corresponds to the intersection of the respective surmise functions. `kmmesh() delivers the mesh of two knowledge structures (Falmagne and Doignon 1998).

kmnotions() computes the notions for a knowledge structure, i.e. the classes of equivalent items. kmeqreduction() gives the structure reduced to these notions. More general is kmsubstructure() which gives the reduction of a knowledge structure to an arbitrary subset of items. kmexpand() is the counterpart to that as it expands a knowledge structure to a larger items set.

Finally, there are functions providing trivial knowledge spaces: kmminimalspace() produces a knowledge space containing just the empty set \(\emptyset\) and the full item set \(Q\). kmmaximalspace(), on the other side, produces the power set \(2^Q\) as knowledge space.

3.2.4 Working with data

There exist several functions for working with data. They can be summarized into three groups.

  1. Generating structures from data: The function kmgenerate() provides a simplistic method to generate a knowledge structure from data by taking all those response patterns as knowledge states which occur with a frequency beyond a certain threshold. This method requires very high numbers of response patterns.

    simdata <- kmsimulate(ksp, 2000, beta=0.1, eta=0.1)
    kgs <- kmgenerate(simdata)
    cat(str_wrap(kmprettyprint(kmbasis(kgs)), width=60))
    #> {{I1}, {I2}, {I3}, {I2, I5}, {I1, I6}, {I2, I6}, {I3, I6},
    #> {I1, I3, I4}, {I3, I5, I6}, {I2, I3, I4, I6}, {I1, I2, I3,
    #> I4, I6, I7}}

    A better approach is given by the DAKS package described in Section 5 below. The function kmiita2SR() generates a surmise relation from an iita object obtained with DAKS functions. Please note, howevr, that DAKS produces a surmise relation and thus a (larger) quasi-ordinal knowledge space.

  2. Validating structures with data: There are two functions implementing different principles of validation: kmvalidate() delivers two distance measures based on the distribution of distances of the response patterns from the structure. kmSRvalidate() validates data against a surmise relation delivering two measures based on how often pairs in the relation are violated by the data.

  3. Generating data through simulation: kmsimulate() simulates response patterns applying the BLIM model.

3.2.5 Knowledge assessment

Adaptive knowledge assessment was the original aim behind developing KST (Doignon and Falmagne 1985). There is quite some literature on this topic (see, e.g., Degreef et al. 1986; Falmagne and Doignon 1988b; Dowling and Hockemeyer 2001; Hockemeyer 2002; Augustin et al. 2015). An assessment routine usually has three key components:

  1. A questioning rule selecting the next test items to be asked,
  2. An update rule determining how to update the user’s profile after evaluating their response to the test item, and
  3. A stopping criterion to decide when enough information has been collected and the assessment is finished.

kstMatrix implements the stochastic procedures introduced by Falmagne and Doignon (1988b) in the notation used by Doignon and Falmagne (1999) (Chapter 10).

For the questioning rule, the halfsplit rule and the informative rule are implemented. For the update rule, the Bayesian and the informative rules are implemented. As a stopping criterion, a probability threshold can be specified; the assessment stops when the maximal likelihood of a knowledge state exceeds this threshold.

There is a general kmassess() function and a simpler kmsassess() version; the former provides for item-specific update parameters while the latter assumes that the update parameters are identical for all items. Furthermore, kmsassess()starts with an equal probability distribution over the knowledge structure while kmassess() provides for more detailed information also here.

We do an assessment with halfsplit questioning rule and Bayesian update rule assuming a general caress error probability of 0.1 and a lucky guess probability of 0.05 stopping at a probability threshold of 0.6.

kmsassess(c(1,1,1,1,0,0,0), ksp, "halfsplit", "Bayesian", 
          0.1, 0.05, NULL, NULL, 0.6)
#> $state
#> [1] 1 1 1 1 0 0 0
#> 
#> $probs
#>  [1] 1.246134e-04 2.243041e-03 2.243041e-03 4.037474e-02
#>  [5] 1.246134e-04 2.243041e-03 2.243041e-03 4.037474e-02
#>  [9] 7.267453e-01 2.361096e-04 4.249972e-03 2.361096e-04
#> [13] 4.249972e-03 7.649950e-02 1.311720e-05 2.361096e-04
#> [17] 2.361096e-04 4.249972e-03 7.649950e-02 2.485364e-05
#> [21] 4.473655e-04 8.052579e-03 8.052579e-03
#> 
#> $queried
#> [1] 6 1 5 2 4
#> 
#> $qtime
#> [1] 0
#> 
#> $utime
#> [1] 0

Five questions (out of 7) have been asked. This is not too much saving but the ratio improves normally considerably with higher item and state numbers. The ninth probability value with 0.727 denotes the highest one.

With an optional additional parameter probdev=TRUE one gets additional information on how the probability distribution has developed over the assessment process. If additionally a directory is specified, this is also illustrated graphically.

Assessment process stepsAssessment process steps

Figure 3.8: Assessment process steps

Figure 9 shows these illustration after the second and after the final, fifth step of the assessment. The differently colored vertices in the graph represent the different likelihoods.

3.2.6 Data sets

kstMatrix provides several data sets, i.e. knowledge spaces. On the one side, there are empirical structures obtained by the group around Cornelia Dowling in the early 90s (Dowling 1993; Baumunk and Dowling 1997) through querying experts.

  • readwrite: Knowledge structures on reading and writing abilities
  • cad: Abilities on working with AutoCAD
  • fractions: Abilities on fractions

Each of these data sets is a list containing different bases.

On the other side, there are two fictional data sets, each a list with different aspects. xpl is a four items structure used four testing purposes in the development of kstMatrix, phsg is a data set used in a publication by Steiner et al. (in press).

3.3 kstIO

The kstIO package offers functions for reading and writing KST files. It support both, set and matrix representations. There are separate reading and writing functions for the various data types.

The reading functions return a list with two elements, matrix and sets if the respective data type is available in the kst package. Otherwise it is only the respective kstMatrix structure.

The classical KST file formats are basically binary text matrices with optional header lines providing additional information. More recently, kstIO has been extended to also support spreadsheet file formats, namely CSV (comma-separated values), ODS (LibreOffice/OpenOffice),, and XLSX (Microsoft Excel).

As an example, we save our basis to an ODS file, re-read it, and look at it within OpenOffice.

write_kbase(phsg$basis,paste0(tempdir(), "/phsg_basis.ods"))
read_kbase(paste0(tempdir(), "/phsg_basis.ods"))
#> $matrix
#>      I1 I2 I3 I4 I5 I6 I7
#> [1,]  1  0  0  0  0  0  0
#> [2,]  0  1  0  0  0  0  0
#> [3,]  0  0  1  0  0  0  0
#> [4,]  1  1  1  1  0  0  0
#> [5,]  0  1  0  0  1  0  0
#> [6,]  0  0  1  0  0  1  0
#> [7,]  1  1  1  1  1  1  1
#> attr(,"class")
#> [1] "kmbasis"  "kmfamset" "matrix"   "array"   
#> 
#> $sets
#> {{"I1"}, {"I2"}, {"I3"}, {"I2", "I5"}, {"I3", "I6"},
#>  {"I1", "I2", "I3", "I4"}, <<set(7)>>}
OpenOffice screenshot of the saved basis file

Figure 3.9: OpenOffice screenshot of the saved basis file

4 KST extensions

4.1 CbKST

In CbKST, we distinguish between the level of observable response behavior to test items and the level of underlying (non-observable) skills and cimpetencies. Within each of these levels, the functions provided by the kstMatrix package (see above) can be used. The CbKST (Hockemeyer 2026a) package caters (only) for the connection between these levels.

Core concepts in CbKST are skill maps and skill multi maps. For the former, there are two interpretations, the conjunctive model and the disjunctive model. Currently, the CbKST package supports only the conjunctive model. For this case, skill maps can be seen (and are treated by the software) as a special case of skill multi maps, i.e. as skill multi maps with exactly one competence per skill.

The starting point for working in CbKST is to read a skill (multi) map from file. This is done with read_skillmultimap(). ODS (OpenOffice/LibreOffice) and XLSX (Microsoft Excel) spreadsheet files are accepted.

With cbkst_perf2comp(), the competence state underlying an observed performance state is determined. cbkst_competencestructure() delivers the set of all competence states underlying some performance state. Specifically for skill maps, there is also a function cbkst_simple_perf2comp().

In the other direction, there are corresponding functions cbkst_comp2perf() and cbkst_performancestructure().

Finally, cbkst_problemfunction() obtains the problem function corresponding to a skill (mult) map.

4.2 pks

pks (Wickelmaier et al. 2024; Heller and Wickelmaier 2013; Brancaccio et al. 2024) offers functions for fitting and testing knowledge structures including parameters for the BLIM and the stochastic learnig model (SLM).

As a preparation, we provide a set of simulated response patterns for our example structure in the NR format used by the pks functions.

simdata <- kmsimulate(ksp, 200, beta=0.1, eta=0.1)
nr <- as.pattern(simdata, freq=TRUE)
head(nr)
#> 0000000 0000001 0000100 0000110 0001000 0010000 
#>       9       1       1       1       2       6

Applying the functions blim() and slm(), we can then estimate the respective models

bm <- blim(ksp, nr)
bm
#> 
#> Basic local independence models (BLIMs)
#> 
#> Number of knowledge states: 23
#> Number of response patterns: 63
#> Number of respondents: 200
#> 
#> Method: Minimum discrepancy
#> Number of iterations: 1
#> Goodness of fit (2 log likelihood ratio):
#>  G2(91) = 102.58, p = 0.19124
#> 
#> Minimum discrepancy distribution (mean = 0.31)
#>   0   1   2   3 
#> 146  47   6   1 
#> 
#> Mean number of errors (total = 0.31)
#> careless error    lucky guess 
#>       0.125000       0.185001 
#> 
#> Error and guessing parameters
#>        beta      eta
#> I1 0.036620 0.000001
#> I2 0.066225 0.000001
#> I3 0.060998 0.000001
#> I4 0.023256 0.070064
#> I5 0.011213 0.043693
#> I6 0.007220 0.032505
#> I7 0.000001 0.086110
kb <- kmbasis(bm$K)
kmprettyprint(kb)
#> [1] "{{I1}, {I2}, {I3}, {I2, I5}, {I3, I6}, {I1, I2, I3, I4}, {I1, I2, I3, I4, I5, I6, I7}}"
kmprettyprint(phsg$basis)
#> [1] "{{I1}, {I2}, {I3}, {I2, I5}, {I3, I6}, {I1, I2, I3, I4}, {I1, I2, I3, I4, I5, I6, I7}}"

We have exemplarily printed the basis of the resulting knowledge structure together with the original basis as comparison, and it shows identical bases. The \(\beta\) and \(\eta\) parameters, however, are partially estimated considerably lower than what we specified in the simulation (0.1 each). This is due to the fact that, in the simulation, by adding noise to some knowledge state, we may well land in some other state.

sm <- slm(ksp, nr)
sm
#> 
#> Simple learning models (SLMs)
#> 
#> Number of knowledge states: 23
#> Number of response patterns: 63
#> Number of respondents: 200
#> 
#> Method: Minimum discrepancy
#> Number of iterations: 1
#> Goodness of fit (2 log likelihood ratio):
#>  G2(106) = 118.18, p = 0.19713
#> 
#> Minimum discrepancy distribution (mean = 0.31)
#>   0   1   2   3 
#> 146  47   6   1 
#> 
#> Mean number of errors (total = 0.31326)
#> careless error    lucky guess 
#>      0.1240278      0.1892363 
#> 
#> Error, guessing, and solvability parameters
#>        beta      eta       g
#> I1 0.036620 0.000001 0.59167
#> I2 0.066225 0.000001 0.75500
#> I3 0.060998 0.000001 0.67625
#> I4 0.023256 0.070064 0.57333
#> I5 0.011213 0.043693 0.54139
#> I6 0.007220 0.032505 0.51201
#> I7 0.000001 0.086110 0.45641
kmprettyprint(kmbasis(sm$K))
#> [1] "{{I1}, {I2}, {I3}, {I2, I5}, {I3, I6}, {I1, I2, I3, I4}, {I1, I2, I3, I4, I5, I6, I7}}"

Here the comments on the BLIM results hold similarly.

4.2.1 Data provided by pks

pks provides several data sets

4.2.2 pksCpp - Probabilistic knowledge structures: additional models and functions

The pksCpp package is the only package presented in this paper that is not available through CRAN. It has to be installed from Github, e.g. by

devtools::install_github("jujupsy/pksCpp")

Primarily, pksCpp (Mollenhauer 2025) re-implements functions from pks in C++. This holds for the conversion functions as.binmat() and as.pattern() as well as for the new function blimCpp().

Furthermore, there is a new functionality by the glimCpp() function which provides a generalized local independence model.

5 Specific functionalities

5.1 CDSS

CDSS (Hockemeyer 2026b) aims at deriving structures from an assignment of taught and required, respectively, skills to learning objects in an existing course. These structures are to be seen as rapid prototypes which need further refinement. Nevertheless, it should drastically simplify the process of developing such structures.

As first step, a skill assignment is read from a respective spreadsheet file using cdss_read_skill_assignment(). This function checks certain conditions and rejects the data by default if these conditions are not met. Afterwards, attribution functions on skill and learning object level can be obtained through cdss_skill_sa2af()and cdss_lo_sa2af(), respectively.

Under certain conditions, a skill assignment may also lead to attribution relations instead of attibution functions. This can be checked with cdss_sa_describes_sr(). If TRUE, the attribution relations can be determined with cdss_skill_sa2ar() and cdss_lo_sa2ar().

Attribution functions and attribution relations can subsequently be closed to surmise functions and surmise relations, respectively, applying kstMatrix::kmclosure(). These can then be further processed usingkstMatrix` (see above) functions.

Furthermore, CDSS offers the cdss_sa2mu() function which creates an interface to the CbKST package (see above)

5.2 DAKS

The DAKS package (Sargin and Ünlü 2016; Ünlü and Sargin 2010) provides functions and a data set for the inductive item tree analysis. Most of its functions are primarily aimed for internal usage. The main function is iita(), its second parameter selects the concrete method to be used.

simdata <- kmsimulate(ksp, 100, beta=0.1, eta=0.1)
i <- iita(simdata, v=1)
i
#> 
#>       Inductive Item Tree Analysis
#> 
#> Algorithm: minimized corrected IITA
#> 
#> quasi order: {(1L, 4L), (1L, 7L), (2L, 1L), (2L, 4L), (2L, 5L),
#>  (2L, 6L), (2L, 7L), (3L, 4L), (3L, 5L), (3L, 6L),
#>  (3L, 7L), (4L, 7L), (5L, 7L), (6L, 4L), (6L, 7L)}
hasse(i$implications, items=7)
Implications determined by `DAKS` compared to the true relationImplications determined by `DAKS` compared to the true relation

Figure 5.1: Implications determined by DAKS compared to the true relation

5.2.1 Data provided by DAKS

DAKS provides a data set from the 2003 PISA (Programme for International Student Assessment) test. It contains 340 response patterns on a 5-item-test on mathematical literacy.

5.3 mycaas

mycaas (Brancaccio and Granziol 2024) provides functions and Shiny apps for the adaptive assessment of knowledge. Actually, the halfsplit questioning rule and the multiplicativ update rule are applkied.

The core routine is the assessment() function. Several parts of the assessment are realized in own functions. mycaas includes two Shiny app functions run_Assessment() and run_Practice(). The result can then be further inspected, e.g. plotting a Hasse diagram of the knowledge structure with the assessed knowledge state and its neighbors being marked.

``mycaas` assessment result

Figure 5.2: `mycaas assessment result

5.3.1 Data

The data set AA_knowledge_test serves as an example for testing the functions of mycaas. It contains a six-items-test on computerized adaptive assessment.

List of described R packages

References

Albert, Dietrich, and Josef Lukas, eds. 1999. Knowledge Spaces: Theories, Empirical Research, Applications. Lawrence Erlbaum Associates. https://doi.org/10.4324/9781410602077.
Augustin, Thomas, Cord Hockemeyer, Reinhard Suck, Patrick Podbregar, Michael D. Kickmeier-Rust, and Dietrich Albert. 2015. “Individualized Skill Assessment in Educational Games: The Mathematical Foundations of Partitioning.” Journal of Mathematical Psychology 67: 1–7. https://doi.org/10.1016/j.jmp.2015.05.003.
Baumunk, Katja, and Cornelia E. Dowling. 1997. “Validity of Spaces for Assessing Knowledge about Fractions.” Journal of Mathematical Psychology 41: 99–105. https://doi.org/10.1006/jmps.1997.1152.
Brancaccio, Andrea, Debora de Chiusole, and Florian Wickelmaier. 2024. “Software Packages for Knowledge Space Theory.” In Knowledge Structures: Recent Developments in Theory and Application, edited by Jürgen Heller and Luca Stefanutti, vol. 7. Advanced Series on Mathematical Psychology. World Scientific. https://doi.org/10.1142/9789811280481_0012.
Brancaccio, Andrea, and Umberto Granziol. 2024. Mycaas: My Computerized Adaptive Assessment. University of Padua, Italy. https://CRAN.R-project.org/package=mycaas.
Chiusole, Debora de, and Luca Stefanutti. 2013. “Modeling Skill Dependence in Probabilistic Competence Structures.” Electronic Notes in Discrete Mathematics 42: 41–48. https://doi.org/10.1016/j.endm.2013.05.144.
Degreef, Eric, Jean-Paul Doignon, André Ducamp, and Jean-Claude Falmagne. 1986. “Languages for the Assessment of Knowledge.” Journal of Mathematical Psychology 30: 243–56. https://doi.org/10.1016/0022-2496/86/90032-5.
Doignon, Jean-Paul. 1994. “Knowledge Spaces and Skill Assignments.” In Contributions to Mathematical Psychology, Psychometrics, and Methodology, edited by Gerhard H. Fischer and Donald Laming. Springer–Verlag. https://doi.org/10.1007/978-1-4612-4308-3_8.
Doignon, Jean-Paul, and Jean-Claude Falmagne. 1985. “Spaces for the Assessment of Knowledge.” International Journal of Man-Machine Studies 23: 175–96. https://doi.org/10.1016/50020-7373(85)80031-6.
Doignon, Jean-Paul, and Jean-Claude Falmagne. 1999. Knowledge Spaces. Springer–Verlag. https://doi.org/10.1007/978-3-642-58625-5.
Dowling, Cornelia E. 1993. “Applying the Basis of a Knowledge Space for Controlling the Questioning of an Expert.” Journal of Mathematical Psychology 37: 21–48. https://doi.org/10.1006/jmps.1993.1002.
Dowling, Cornelia E., and Cord Hockemeyer. 2001. “Automata for the Assessment of Knowledge.” IEEE Transactions on Knowledge and Data Engineering 13 (3): 451–61. https://doi.org/10.1109/69.929902.
Düntsch, Ivo, and Günther Gediga. 1995. “Skills and Knowledge Structures.” British Journal of Mathematical and Statistical Psychology 48: 9–27. https://doi.org/10.1111/j2044-8317.1995.tb01047.x.
Falmagne, Jean-Claude, and Jean-Paul Doignon. 1988a. “A Class of Stochastic Procedures for the Assessment of Knowledge.” British Journal of Mathematical and Statistical Psychology 41: 1–23. https://doi.org/10.1111/2044-8317.1988.tb00884.x.
Falmagne, Jean-Claude, and Jean-Paul Doignon. 1988b. “A Markovian Procedure for Assessing the State of a System.” Journal of Mathematical Psychology 32: 232–58. https://doi.org/10.1016/0022-2496(88)90011-9.
Falmagne, Jean-Claude, and Jean-Paul Doignon. 1998. “Meshing Knowledge Structures.” In Recent Progress in Mathematical Psychology, edited by Cornelia E. Dowling, Fred S. Roberts, and Peter Theuns. Scientific Psychology Series. Lawrence Erlbaum Associates Ltd.
Heller, Jürgen, and Luca Stefanutti, eds. 2024. Knowledge Structures: Recent Developments in Theory and Application. Vol. 7. Advanced Series on Mathematical Psychology. World Scientific. https://doi.org/10.1142/13519.
Heller, Jürgen, and Florian Wickelmaier. 2013. “Minimum Discrepancy Estimation in Probabilistic Knowledge Structures.” Electronic Notes in Discrete Mathematics 42: 49–56. https://doi.org/10.1016/j.endm.2013.05.145.
Hockemeyer, Cord. 2002. “A Comparison of Non–Deterministic Procedures for the Adaptive Assessment of Knowledge.” Psychologische Beiträge 44: 495–503.
Hockemeyer, Cord. 2026a. CbKST: Functions for Competence–Based Knowledge Space Theory. Department of Psychology, University of Graz, Austria. https://doi.org/10.32614/CRAN.package.CbKST.
Hockemeyer, Cord. 2026b. CDSS: Course–Dependent Skill Structures. Department of Psychology, University of Graz, Austria. https://doi.org/10.32614/CRAN.package.CDSS.
Hockemeyer, Cord, Peter Steiner, and Wai Wong. 2026. KstMatrix: Basic Functions in Knowledge Space Theory Using Matrix Representations. Department of Psychology, University of Graz, Austria. https://doi.org/10.32614/CRAN.package.kstMatrix.
Korossy, Klaus. 1997. “Extending the Theory of Knowledge Spaces: A Competence–Performance Approach.” Zeitschrift für Psychologie 205: 53–82.
Mollenhauer, Julian. 2025. PksCpp: Probabilistic Knowledge Structures — Additional Models and Functions. University of Tübingen, Germany. https://github.com/jujupsy/pksCpp/.
Sargin, Anatol, and Ali Ünlü. 2016. DAKS: Data Analysis and Knowledge Spaces. Comprehensive R Archive Network. https://doi.org/10.32614/CRAN.package.DAKS.
Schrepp, Martin, Theo Held, and Dietrich Albert. 1999. “Component–Based Construction of Surmise Relations for Chess Problems.” In Knowledge Spaces: Theories, Empirical Research, Applications, edited by Dietrich Albert and Josef Lukas. Lawrence Erlbaum Associates. https://doi.org/10.4324/9781410602077.
Stahl, Christina. 2008. “Developing a Framework for Competence Assessment.” Unpublished Dissertation, Vienna University of Economics; Business.
Stahl, Christina, David Meyer, and Cord Hockemeyer. 2026. Kst: Knowledge Space Theory. Department of Psychology, University of Graz, Austria. https://doi.org/10.32614/CRAN.package.kst.
Stefanutti, Luca, and Debora de Chiusole. 2017. “On the Assessment of Learning in Competence Based Knowledge Space Theory.” Journal of Mathematical Psychology 80: 22–32. https://doi.org/10.1016/j.j,p.2017.008.003.
Steiner, Peter, Cord Hockemeyer, Miachael Kickmeier–Rust, and Jan Hochweber. in press. “Structuring and Assessing Knowledge with Knowledge Space Theory.” In Quantitative Methods in Educational Research:: Concepts and Applications, edited by M. L. Wilson, A. A. Tawfik, A. D. Ritzhaupt, and A. T. Lowdermilk. EdTech Books. https://edtechbooks.org/qmerca/knowledgespacetheory.
Taagepera, Mare, Frank Potter, George E. Miller, and Kamakshi Lakshminarayan. 1997. “Mapping Students’ Thinking Patterns by the Use of Knowledge Space Theory.” International Journal of Science Education 19: 283–302. https://doi.org/10.1080/0950069970190303.
Ünlü, Ali, and Anatol Sargin. 2010. “DAKS: An R Package for Data Analysis Methods in Knowledge Space Theory.” Journal of Statistical Software 37: 1–31. https://doi.org/10.18637/jss.v037.i02.
Wickelmaier, Florian, Jürgen Heller, Julian Mollenhauer, and Pasquale Anselmi. 2024. Pks: Probabilistic Knowledge Structures. University of Tübingen, Germany. https://doi.org/10.32614/CRAN.package.pks.