Using it¶
The shortest way to start¶
Every example takes the path to your data and prints what it found. There is nothing to set up first and no model to train.
A .sym file has one symbol per line. Any sequence works: characters, words, byte values, note numbers, residue names. The measure never learns what a symbol means, and it can't tell one subject from another.
If you run an example with no argument, it prints how to use it and stops.
The six parts¶
| part | what you call it for |
|---|---|
representation |
turn your object into points carrying values |
partition |
choose the unit and the scale to read at |
reference |
build the maximum entropy background the object's own constraints allow |
measure |
read the departure from that background |
sift |
filter candidates with a necessary condition |
oracle |
check against ground truth somebody else published |
Only representation knows what kind of thing your object is. Under it are atom, constants, game, particle, picture, sound, structure and text. Everything after it sees only points and values.
archive/src/python/README.md is the map.
Where everything else is¶
| what it operates on | |
|---|---|
src/ |
points and values, no domain. The engine |
evidence/ |
the claims. The proofs, and the R and MATLAB ports |
examples/ |
a corpus, through src/. Numbered demonstrations |
utils/test/ |
the engine. The correctness checks in utils/test/src/, laid out as src/ is, and the published test vectors |
utils/bench/ |
the engine. The benches: how fast it is |
theory/ |
the research papers, and the ledger they cite |
utils/maint/ |
the repository itself. Records, gates, prose checks, fetchers, the research paper build |
utils/maint/ is sorted into categories, with no loose scripts. utils/maint/README.md says what goes in each one. For example, utils/maint/data/ holds outside material, and utils/maint/analysis/ holds the surveys the research papers ask for.
Reading the result¶
Every measurement comes with a floor. The floor is what the same measure gives on a shuffled copy of the same symbols, a version with no structure at all.
A reading below its floor found nothing. That's why every example prints the floor next to the number.
The floor changes with the size of the sample. A floor worked out on a large body of text tells you nothing about a short file. The examples work out the floor at the size actually measured.
The twelve subjects¶
examples/ runs the same six parts from start to finish on real data. The proofs behind the ledger are kept separately, in evidence/proofs/posits/.
| domain | examples | what it reads |
|---|---|---|
language |
60 | corpora, orthographies, dialect borders |
any_corpus |
19 | any symbol sequence, domain unspecified |
game_theory |
16 | games, by how open the result stays after a move |
proteins |
13 | backbone coordinates |
particle_physics |
11 | atoms and particles as exact quantum numbers |
crystallography |
9 | cell edges, against published ones |
art |
9 | images as byte sequences |
cell_tracking |
7 | cell positions across frames |
chemistry |
5 | molecules as atoms and bonds |
source |
4 | source code as a symbol stream |
sound |
3 | recordings as bit fields |
molecules |
3 | molecular formulae, legal from illegal by valence |
Every example has a catalog number at the top of the file, such as LNG-4-012. If you cite that number, the citation still works after the file moves. utils/maint/catalog/catalog.py keeps the list.
The search kernel¶
anchor_steer_count counts how many times a needle (the pattern you search for) appears in a corpus (the text you search in). Its last argument is 1 to order the probes by rarity, or 0 to keep them in the order they appear. The count is the same either way (orior_descent.h:342-368). The function borrows both buffers for the length of the call ([BORROWS]).
#include <stdint.h>
#include <stdio.h>
#include <string.h>
#include "orior.h"
int main(void)
{
const char *corpus = "abracadabra abracadabra";
const char *needle = "abra";
const size_t corpus_len = strlen(corpus);
const size_t needle_len = strlen(needle);
const size_t count = anchor_steer_count((const uint8_t *)corpus, corpus_len,
(const uint8_t *)needle, needle_len, 1);
printf("%lu occurrences, scanned by the %s engine\n", (unsigned long)count,
anchor_steer_best_engine()->name);
return 0;
}
Built with the gcc line in Setup, under gcc on x86-64 Windows, it prints 4 occurrences, scanned by the portable engine.
The sift is a safe filter: no arrangement of anchors can lose a true occurrence.
- It can only be wrong in one direction. Any mistake is an over-count, never a missed match.
- Its memory is small. It keeps
mbits of state for a pattern of lengthm, with no table over the alphabet. An alphabet of real numbers, or one too large to list, costs it nothing extra.
The Python in archive/src/python/engine/nbody/orior/sift/ builds the same thing and shares no code with the C version. The two are checked against each other by making sure their counts agree.
Other languages¶
| language | file | status |
|---|---|---|
| R | evidence/sims/r/departure.R |
runs, checked against the reference |
| MATLAB and Octave | evidence/sims/matlab/orior_departure.m |
run on Octave 11.3.0, inside the reference floor; MATLAB proper not run here |
Each language draws its shuffle from its own random number generator, and no two agree to the last digit. A port counts as correct when its result lands inside the reseeding floor of the Python version.
This was checked on 200000 symbols over twelve seeds:
- A clustered sequence reads 0.4228 in Python and 0.4282 in R, against a floor of 0.0092.
- A memoryless sequence reads 0.9953 and 0.9933, against a floor of 0.0044.
Both gaps come to about half a floor.
If you are working on a language¶
Read the conditions of use first. These tools can generate language. Output near the edge of what the source covers can read smoothly and still not be the language. Nothing here marks which side of that line a result fell on. A person must review the output. That review is a condition of use.
Author: dstroy0 (Douglas Quigg) dquigg123@gmail.com