raispace
← back to blog

Epistemic Physics 1: Bayesian Special Relativity

· 45 min epistemic-physics
wWas aiAI uUsed hHere?

nNo aiAI was used for the writing and structure of this article, nor for the primary ideas behind it. iI used aiAI in three ways - guiding it to create interactive demos the way iI wanted, double checking work, and making more connections with fields iI'm less familiar with.

cConnecting eEpistemology, sStatistics, pPhysics, and mMore

iIn this series on epistemic physics, iI'll show you how treating probability as relativistic velocity and evidence as spacetime leads to all sorts of wonderful connections. hHere are some highlights!

𝐸[𝐷𝑠]=21𝑠(𝑠1)Γ(𝑠+1)𝜁(𝑠),𝐷=|Φ1Φ2|=|Φ1+Φ2|
$$ E[D^s] = 2^{1-s}(s-1)\Gamma(s+1)\zeta(s), \qquad D=|\Phi_1-\Phi_2|=|\Phi_1+\Phi_2| $$
  • iIn addition, multiplying uniform beliefs together can result in "increasing" zetas (increasing n) too:
𝐸[11𝐵1𝐵2]=𝜁(2)=𝜋26
$$ E\left[\frac1{1-B_1B_2}\right] = \zeta(2) = \frac{\pi^2}{6} $$
𝐸[11𝑛𝑖=1𝐵𝑖]=𝜁(𝑛),𝑛>1
$$ E\left[ \frac1{1-\prod_{i=1}^nB_i} \right] = \zeta(n), \qquad n>1 $$
  • tThis is the waiting time until two-coin flips come up with two heads. hHave fun mathematicians (and aiAI agents iI guess)!
    • iI hope to expand on how epistemic physics relates to the rRiemann hHypothesis, if it doesn't get solved by then, in a future post.

aAnd check out the fFun pPredictions section below! iI won't prove all these connections in this first introductory article, but feel free to explore yourself! lLuckily, the essentials of epistemic physics are extremely easy to understand with some basic knowledge of statistics and relativity.

aAbout my other websites…

aAs of 2026-09-15:

https://epistemicphysics.com is my aiAI slop website on epistemic physics, iI use it as a dumping ground for ideas with little curation; almost all of it is aiAI generated and rather difficult to understand. eEnter at your own risk. iI may de-slopify it someday if iI have the time.

https://sloth.ink is currently heavily outdated - iI am in the process of overhauling the underlying mathematics in light of epistemic physics - but the articles may still provide interesting insights on beliefs on propositions.

eEssentials

bBeta dDistribution with pPrior

dDefine a belief 𝐵$B$ distributed according to a bBeta distribution

𝐵Beta(𝛼,𝛽),𝛼,𝛽>0.
$$ B\sim\operatorname{Beta}(\alpha,\beta), \qquad \alpha,\beta>0. $$

with mean

𝜇=𝐸[𝐵]=𝛼𝛼+𝛽
$$ \mu = E[B] = \frac{\alpha}{\alpha + \beta} $$

dDefine the total evidence and signed evidence balance by

𝜏E=𝛼+𝛽,𝑥E=𝛼𝛽.
$$ \tau_{\mathrm E}=\alpha+\beta, \qquad x_{\mathrm E}=\alpha-\beta. $$

hHere, upright subscript E$\mathrm{E}$ stands for "epistemic." nNow, linearly recenter the distribution between [1,1]$[-1,1]$. dDefine the centered random variable

𝐵E=2𝐵1
$$ B_{\mathrm{E}} = 2B-1 $$

tThe recentered mean is thus:

𝑣E=𝔼[𝐵E]=2𝜇1=2𝛼𝛼+𝛽1=2𝛼(𝛼+𝛽)𝛼+𝛽=𝛼𝛽𝛼+𝛽=𝑥E𝜏E
$$ \begin{aligned} v_\mathrm{E} &= \mathbb E[B_{\mathrm{E}}] \\ &= 2\mu-1 \\ &=2\frac{\alpha}{\alpha+\beta}-1 \\ &=\frac{2\alpha-(\alpha+\beta)}{\alpha+\beta} \\ &=\frac{\alpha-\beta}{\alpha+\beta} \\ &= \frac{x_{\mathrm{E}}}{\tau_{\mathrm{E}}}\\ \end{aligned} $$

Beta vs Velocity

The same Beta distribution, original vs recentered.

density over probability μ same belief over velocity vE — twice as wide, half as tall
1234μ0½1vE−10+1mean · drag me
each observation adds one full unit of evidence

α = 0.50 β = 0.50 τE = 1.00 μ = 0.500 vE = 0.000

cCones

lLight cCones

fFor a 1+1$1+1$ dimensional light cone,

𝑑𝑠2=𝑐2𝑑𝑡2𝑑𝑥2
$$ ds^2 = c^2dt^2 - dx^2 $$

where 𝑠$s$ is the spacetime interval, 𝑡$t$ is time, and 𝑥$x$ is position. aAt the origin,

𝑠2=𝑐2𝑡2𝑥2
$$ s^2 = c^2t^2 - x^2 $$

tThe light-cone coordinates, or null coordinates are:

𝑥+=𝑐𝑡+𝑥,𝑥=𝑐𝑡𝑥
$$ x^+= ct+x, \qquad x^-=ct-x $$

eEvidence cones

rRecall the mean of our recentered bBeta distribution:

𝑣E=𝛼𝛽𝛼+𝛽=𝑥E𝜏E
$$ v_\mathrm{E} =\frac{\alpha-\beta}{\alpha+\beta} = \frac{x_{\mathrm{E}}}{\tau_{\mathrm{E}}} $$

tThe prior-inclusive bBeta parameters are proportional to the null coordinates of a light cone:

𝜏E=𝛼+𝛽,𝑥E=𝛼𝛽.
$$ \tau_{\mathrm E}=\alpha+\beta, \qquad x_{\mathrm E}=\alpha-\beta. $$

hHowever, the inverse transformation also looks like the null coordinates of a light cone:

𝛼=𝜏E+𝑥E2,𝛽=𝜏E𝑥E2.
$$ \alpha=\frac{\tau_{\mathrm E}+x_{\mathrm E}}{2}, \qquad \beta=\frac{\tau_{\mathrm E}-x_{\mathrm E}}{2}. $$

wWhich representation do we choose? nNote that the inverse relations are derived by

𝜏E+𝑥E=(𝛼+𝛽)+(𝛼𝛽)=2𝛼𝜏E𝑥E=(𝛼+𝛽)(𝛼𝛽)=2𝛽
$$ \begin{aligned} \tau_{\mathrm E}+x_{\mathrm E} &=(\alpha+\beta)+(\alpha-\beta) \\ &=2\alpha \\ \tau_{\mathrm E}-x_{\mathrm E} &=(\alpha+\beta)-(\alpha-\beta) \\ &=2\beta \end{aligned} $$

tThese both look like conic coordinates, and both possible coordinate constructions result in valid geometry. hHowever, only one configuration preserves beta evidence semantics.

pPropertysSemantics-preserving geometryaAlternate geometry
cCone coordinates(𝜏E,𝑥E)$(\tau_{\mathrm E},x_{\mathrm E})$(𝛼,𝛽)$(\alpha,\beta)$
nNull coordinates𝛼,𝛽$\alpha,\beta$𝜏E,𝑥E$\tau_{\mathrm E},x_{\mathrm E}$
mMetric𝜏2E𝑥2E=4𝛼𝛽$\tau_{\mathrm E}^2-x_{\mathrm E}^2=4\alpha\beta$𝛼2𝛽2=𝜏E𝑥E$\alpha^2-\beta^2=\tau_{\mathrm E}x_{\mathrm E}$
nNull boundary𝛼=0$\alpha=0$ or 𝛽=0$\beta=0$𝜏E=0$\tau_{\mathrm E}=0$ or 𝑥E=0$x_{\mathrm E}=0$
cCausal sector𝛼,𝛽0$\alpha,\beta\ge 0$, equivalently 𝜏E𝑥eE
bBalanced evidence𝑥E=0$x_{\mathrm E}=0$, interior centerline𝛼=𝛽$\alpha=\beta$, hence 𝑥E=0$x_{\mathrm E}=0$, a null boundary
cComplement 𝛼𝛽$\alpha\leftrightarrow\beta$pPreserves the intervalrReverses the sign of the interval

iIn the alternate geometry, 𝛼$\alpha$ is timelike and 𝛽$\beta$ is spacelike (you could also define 𝑥E=𝛽𝛼$x_{\mathrm{E}}=\beta-\alpha$ to swap those). eEach success outcome, or event observation, results in time progression, while each failure outcome, or null observation, results in space progression. oOnly probabilities where 𝑝>0.5$p>0.5$ would be causal, and failure-dominant probabilities become spacelike. tThis seems arbitrary, privileging one evidential outcome. iIn addition, balanced evidence, or conflict, becomes indistinguishable from no evidence, or vacuity, since 𝛼2𝛽2=0$\alpha^2-\beta^2=0$.

oOn the other hand, if we set 𝛼$\alpha$ and 𝛽$\beta$ to be the null coordinates, we gain a coherent interpretation. hHere, 𝜏E$\tau_{\mathrm{E}}$, interpreted as total observations, is timelike, and signed evidence balance 𝑥E$x_{\mathrm{E}}$, interpreted as net support, is spacelike (alternatively you could use net opposition, but net support is more intuitive). tThe epistemic velocity 𝑣E=𝑥E/𝜏E$v_{\mathrm{E}}=x_{\mathrm{E}}/\tau_{\mathrm{E}}$ becomes recentered on 0$0$ between 1$-1$ and 1$1$. oOver time you gain evidence, while a change in net support moves you. cConflict is distinguishable from vacuity, since more conflict moves you further in the timelike direction. fFor a proper bBeta distribution, 𝛼,𝛽>0$\alpha,\beta>0$, and therefore every proper bBeta state lies strictly inside the future lLorentz cone.

hHow iI got here

oOn aAugust 12, 2026, while exploring the concepts of negative and complex evidence in the context of stochastic valence networks (neural networks using bBeta distributions and mixtures), iI noticed the connection between special relativity and bBeta-distributed evidence. iI've been thinking about evidence, bBetas and dDirichlets for a while now, thanks to my other project, gGated eEpistemic cCalculus and the currently outdated sSlothink website, which aim to reduce misinformation and encourage empathy by helping people make their subjective beliefs consistent and free of hypocrisy. iIn valence networks, membrane potential of valence networks is analogous to position, the current is analogous to time, and the mean firing rate is analogous to velocity.

eEnergy-momentum cCones

tThe energy-momentum relation is

𝐸2=(𝑝𝑐)2+(𝑚𝑐2)2(𝐸𝑐)2𝑝2=𝑚2𝑐2
$$ \begin{aligned} E^2 &= (pc)^2 + (mc^2)^2 \\ \left(\frac{E}{c}\right)^2 - p^2 &= m^2c^2 \\ \end{aligned} $$

tThe corresponding null coordinates are

𝑝+=𝐸𝑐+𝑝,𝑝=𝐸𝑐𝑝
$$ p^+ = \frac{E}{c}+p, \qquad p^- = \frac{E}{c}-p $$

eEnergy-momentum space thus shares the same lLorentz-cone geometry as spacetime and the evidence cone. tTo compare energy-momentum space to evidence space, use natural units with 𝑐=1$c=1$:

𝑚2=𝐸2𝑝2=𝜏2E𝑥2E=(𝛼+𝛽)2(𝛼𝛽)2=𝛼2+2𝛼𝛽+𝛽2𝛼2+2𝛼𝛽𝛽2=4𝛼𝛽=𝑚2E
$$ \begin{aligned} m^2&=E^2-p^2\\ &=\tau_{\mathrm{E}}^2-x_{\mathrm{E}}^2\\ &=(\alpha + \beta)^2 - (\alpha - \beta)^2 \\ &=\alpha^2 + 2\alpha\beta + \beta^2 - \alpha^2 + 2\alpha\beta - \beta^2 \\ &=4\alpha\beta \\ &=m_{\mathrm{E}}^2 \end{aligned} $$

aAs an analogy, one can interpret 𝜏𝐸$\tau_{E}$ as energy-like, 𝑥E$x_{\mathrm{E}}$ as momentum-like, 𝑚E=2𝛼𝛽$m_{\mathrm{E}} = 2\sqrt{ \alpha\beta }$ as mass-like, and 𝑣E$v_{\mathrm{E}}$ still as epistemic velocity. hHowever, in physics, energy-momentum describes interactions, such as the change in momentum in a collision. oOther differences from spacetime include additivity, where the total momentum of system is the sum of parts, and tangency, where momentum is tangent to the worldline while position is a point on the worldline. tThus the better analogy is to consider incremental changes in evidence:

𝑑𝜏E=𝑑𝛼+𝑑𝛽,𝑑𝑥E=𝑑𝛼𝑑𝛽
$$ \begin{aligned} d\tau_{\mathrm{E}} = d \alpha +d \beta, \qquad dx_{\mathrm{E}}=d \alpha -d \beta \end{aligned} $$

tThis preserves the notion that energy and momentum are conserved and exchanged during interactions.

vVariance

nNow we can derive some identities for the variance of a recentered belief:

Var(𝐵E)=4𝛼𝛽(𝛼+𝛽)2(𝛼+𝛽+1)=1𝑣2E𝜏E+1=𝑚2E𝜏2E(𝜏E+1)
$$ \begin{aligned} \operatorname{Var}(B_{\mathrm E}) &=\frac{4\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}\\ &=\frac{1-v_{\mathrm E}^2}{\tau_{\mathrm E}+1}\\ &=\frac{m_{\mathrm E}^2}{\tau_{\mathrm E}^2(\tau_{\mathrm E}+1)}\\ \end{aligned} $$

Two cones, One geometry

Evidence spacetime: the worldline
all oppositionall supportτExE
Energy-momentum
β photonsα photonsE-likep-likeαβ
or drag either dot

Y = 0 N = 0 τE = 1.0 xE = 0.0 vE = 0.000 mE = 1.00 γE = 1.000

$$m_{\mathrm E}^2=\tau_{\mathrm E}^2-x_{\mathrm E}^2=4\alpha\beta$$

bBeyond the essentials

eEvidence nNotation

iIt is often useful to separate raw evidence from the prior. dDefine a neutral symmetric prior concentration 𝑛0$n_0$

𝛼0=𝛽0=𝑛02,𝐵Beta(𝑛02,𝑛02)
$$ \alpha_{0}=\beta_{0}=\frac{n_{0}}{2}, \qquad B\sim \operatorname{Beta}\left( \frac{n_{0}}{2},\frac{n_{0}}{2} \right) $$

dDefine the total evidence accumulated over the prior 𝑡E$t_\mathrm{E}$ as trials 𝑌$Y$, for support, and 𝑁$N$, for opposition:

𝑡E=𝑌+𝑁
$$ t_\mathrm{E} = Y+N $$

wWhile bBeta evidence trials may be continuous, one can interpret integer-valued trials as bBernoulli outcomes. oOur previous notation can thus be defined like:

𝛼=𝑌+𝑛02,𝛽=𝑁+𝑛02𝜏E=𝑌+𝑁+𝑛0𝑥E=𝑌+𝑛02𝑁𝑛02=𝑌𝑁𝑣E=𝑌𝑁𝑌+𝑁+𝑛0=𝑥E𝑡𝐸+𝑛0
$$ \begin{aligned} \alpha &= Y+\frac{n_{0}}{2}, \qquad \beta=N+\frac{n_{0}}{2}\\ \tau_{\mathrm{E}} &= Y+N+n_{0} \\ x_{\mathrm{E}} &= Y+\frac{n_{0}}{2} - N - \frac{n_{0}}{2}\\ &=Y-N \\ v_{\mathrm{E}} &= \frac{Y-N}{Y+N+n_{0}} \\ &=\frac{x_{\mathrm{E}}}{t_{E} + n_{0}} \end{aligned} $$

tThe jJeffreys pPrior

tThe jJeffreys prior, denoted as 𝜛$\varpi$ below, or the arcsine distribution, is a particularly suitable prior for epistemic physics.

𝛼=𝑌+12,𝛽=𝑁+12𝜏E=𝑌+𝑁+1𝑣E=𝑌𝑁𝑌+𝑁+1𝜛=Beta(12,12)=1𝜋𝜇(1𝜇)𝜛E=1𝜋1𝑣2E
$$ \begin{aligned} \alpha&=Y+\frac12, \qquad \beta=N+\frac12 \\ \tau_{\mathrm E}&=Y+N+1\\ v_{\mathrm E}&=\frac{Y-N}{Y+N+1}\\ \varpi &= \operatorname{Beta}\left( \frac{1}{2}, \frac{1}{2} \right) = \frac{1}{\pi \sqrt{ \mu(1-\mu) }} \\ \varpi_{\mathrm{E}} &= \frac{1}{\pi \sqrt{ 1-v_{\mathrm{E}}^2 }} \end{aligned} $$

tThe reasoning is as follows. fFirst, using a jJeffreys prior assigns equal prior probability to equal lengths in fFisher information. gGiven 1000 coin flips, changing the heads probability from 49% to 50% changes the expected heads count from 490 to 500. cChanging from 1% to 2% goes from 10 heads to 20 heads. aAlthough both situations add 10 heads, the scales are different. tThe former requires only 1.02 times the heads rate, while the latter requires doubling the rate of heads. fFisher information states:

𝐼(𝜇)=1𝜇(1𝜇)
$$ I(\mu) = \frac{1}{\mu(1-\mu)} $$

tThe fFisher length element, is

𝑑=𝐼(𝜇)|𝑑𝜇|=|𝑑𝜇|𝜇(1𝜇)
$$ d\ell=\sqrt{I(\mu)}\,|d\mu| =\frac{|d\mu|}{\sqrt{\mu(1-\mu)}} $$

aA short interval near probability 0 or 1 covers more fFisher length than an equally wide interval near the center. cCompare the scales - 0 to 1 for probability versus 0 to 𝜋$\pi$ for fFisher length. fFisher length allows for a fixed interval along the entire distance from 0 to 𝜋$\pi$, while the interval varies on the probability scale.

sSecond, consider the identity 𝜏2E𝑥2E=𝑚2E$\tau_{\mathrm E}^2 - x_{\mathrm E}^2 = m_{\mathrm E}^2$ , which can be rearranged a la pPythagoras:

𝜏2E=𝑥2E+𝑚2E
$$ \tau_{\mathrm E}^2 = x_{\mathrm E}^2 + m_{\mathrm E}^2 $$

eEpistemic position and mass lie on a circle of radius 𝜏E$\tau_\mathrm{E}$ (imagine position on the x-axis and mass on the y-axis). sSince epistemic mass is positive for proper beliefs, the upper semicircle contains the normally reachable states.

𝑥E=𝜏Esin𝜃,𝑚E=𝜏Ecos𝜃,𝜋2<𝜃<𝜋2.
$$ x_{\mathrm E} = \tau_{\mathrm E}\sin \theta, \qquad m_{\mathrm E} = \tau_{\mathrm E}\cos \theta, \qquad -\frac{\pi}{2} < \theta < \frac{\pi}{2}. $$

iIn this construction, the uniform, or lLaplace prior, is uniform only along the x-axis, while the jJeffreys prior is uniform along 𝜃$\theta$, and thus uniform along the semicircle. tThe jJeffreys prior has no preference for how evidence is split between mass and position. tThis construction also clearly illustrates that the 𝜋$\pi$ that appears is the angular distance of the semicircle's arc.

tThird, when we extend from 1+1 dD to higher dimensions, the natural extension to the jJeffreys prior works well. sSee the multidimensional extension section for more details.

Jeffreys vs Uniform

−10+1vEμ: 0 on the left, ½ in the middle, 1 on the rightxEmEτEθ = 40°

vE = 0.643 μ = 0.821 mEE = 0.766 γE = 1.305 ϖE = 0.416

$$x_{\mathrm E}=\tau_{\mathrm E}\sin\theta,\qquad m_{\mathrm E}=\tau_{\mathrm E}\cos\theta,\qquad \varpi_{\mathrm E}=\frac{1}{\pi\cos\theta}=\frac{\gamma_{\mathrm E}}{\pi}$$

rRapidity

wWe can relate epistemic velocity to rapidity. dDefine epistemic rapidity as

𝜙E=artanh𝑣E=12log1+𝑣E1𝑣E=12log1+𝛼𝛽𝛼+𝛽1𝛼𝛽𝛼+𝛽=12log𝛼+𝛽+𝛼𝛽𝛼+𝛽𝛼+𝛽𝛼+𝛽𝛼+𝛽=12log2𝛼2𝛽=12log𝛼𝛽
$$ \begin{aligned} \phi_{\mathrm E}&=\operatorname{artanh}v_{\mathrm E} \\ &=\frac{1}{2}\log\frac{1+v_{\mathrm{E}}}{1-v_{\mathrm{E}}} \\ &= \frac12\log \frac{ 1+\frac{\alpha-\beta}{\alpha+\beta} }{ 1-\frac{\alpha-\beta}{\alpha+\beta} }\\ &= \frac12\log \frac{ \frac{\alpha+\beta+\alpha-\beta}{\alpha+\beta} }{ \frac{\alpha+\beta-\alpha+\beta}{\alpha+\beta} }\\ ​&= \frac12\log \frac{ 2\alpha }{ 2\beta} \\ &=\frac{1}{2}\log\frac{\alpha}{\beta} \\ \end{aligned} $$

tThis is half the log-odds. nNote that rapidity is related to bBondi's k-factor, thus forming the connection between the square root of the likelihood ratio and the dDoppler effect. oOne can also define:

𝜏E=𝑚Ecosh𝜙E=𝛼+𝛽𝑥E=𝑚Esinh𝜙E=𝛼𝛽
$$ \begin{aligned} \tau_{\mathrm{E}}&=m_{\mathrm{E}} \cosh \phi_{\mathrm{E}} = \alpha+\beta\\ x_{\mathrm{E}}&=m_{\mathrm{E}}\sinh \phi_{\mathrm{E}}=\alpha-\beta\\ \end{aligned} $$

wWe can go a step further and transform the entire belief, not only the mean:

ΦE=artanh(𝐵E)=12log1+𝐵E1𝐵E=12log𝐵1𝐵
$$ \Phi_{\mathrm E}=\operatorname{artanh}(B_{\mathrm E}) =\frac12\log\frac{1+B_{\mathrm E}}{1-B_{\mathrm E}} =\frac12\log\frac{B}{1-B} $$

nNote that this distinction between translating the velocity versus the whole belief is very important - some connections from earlier only work for one transformation and not the other. tThe inverse transformation is:

𝐵E=tanhΦE,𝐵=1+tanhΦE2.
$$ B_{\mathrm E}=\tanh\Phi_{\mathrm E},\qquad B=\frac{1+\tanh\Phi_{\mathrm E}}{2}. $$

nNote that

𝔼[𝐵E]=𝔼[tanhΦE]=𝑣E=tanh𝜙E,𝔼[ΦE]𝜙E
$$ \mathbb E[B_{\mathrm{E}}] = \mathbb E[\tanh\Phi_{\mathrm E}] =v_{\mathrm E}=\tanh\phi_{\mathrm E},\qquad \mathbb E[\Phi_{\mathrm E}]\ne\phi_{\mathrm E} $$

mMore specifically,

𝐸[ΦE]=12(𝜓(𝛼)𝜓(𝛽)),VarΦE=14(𝜓(𝛼)+𝜓(𝛽))
$$ E[\Phi_{\mathrm{E}}]=\tfrac12(\psi(\alpha)-\psi(\beta)), \qquad \operatorname{Var}\Phi_{\mathrm{E}}=\tfrac14(\psi'(\alpha)+\psi'(\beta)) $$

Rapidity across fields

0½1-4-20+2+4probability μrapidity φE

φE = 0.550 μ = 75.0% vE = 0.501

special relativity

$v/c=\tanh\phi_{\mathrm E}$

logistic regression · softmax · LLR

$\ell=\log\tfrac{\mu}{1-\mu}=2\phi_{\mathrm E}$

chess rating

$\Delta_{\mathrm{Elo}}=\tfrac{800}{\ln 10}\,\phi_{\mathrm E}$

acid-base chemistry

$\mathrm{pH}-\mathrm{p}K_a=\tfrac{2}{\ln 10}\,\phi_{\mathrm E}$

prediction market

$\text{price}=100\,\mu$

psychometrics (Rasch)

$\theta_{\text{ability}}-b_{\text{difficulty}}=2\phi_{\mathrm E}$

ion-channel gating

$V-V_{1/2}=2k\,\phi_{\mathrm E},\ k=6\,\mathrm{mV}$

Nernst equilibrium

$E=\tfrac{2RT}{zF}\,\phi_{\mathrm E}$

ligand binding (Hill, n = 2)

$[L]/K=e^{2\phi_{\mathrm E}/n}$

Fermi-Dirac occupancy

$(E_F-\varepsilon)/k_BT=2\phi_{\mathrm E}$

spin magnetization

$\bar s=\tanh\!\left(h/k_BT\right),\ h/k_BT=\phi_{\mathrm E}$

Kelly betting

$f^{*}=\tanh\phi_{\mathrm E}$

natural selection

$\phi_{\mathrm E}(t)=\phi_0+\tfrac{s}{2}\,t,\ s=5\%$

eEpistemic iInterpretations

iI'm making a deliberate simplification in choosing to model fuzzy beliefs on propositions as bBeta distributions. iIn reality, beliefs may also include additional properties such as how underspecified the proposition is (naturally leading to a nested bBeta or dDirichlet interpretation), or encompass more than propositions (such as numerical estimates, though those can arguably be reduced to propositions). hHowever, the simplification leads to a useful epistemic world with constraints and physical and statistical interpretations. fFor example, one can interpret approaching the speed of light as approaching absolute certainty of belief.

tThe sSpeed lLimit of bBelief

eEpistemically, the speed of light plays several roles. oOne interpretation says that a belief with mass can never reach absolute certainty. aAnother interpretation says that once you see evidence, you can't unsee it. wWe can try a fun thought experiment - what happens epistemically or probabilistically when we try breaking the speed limit? wWe get probabilities less than 0 or greater than 1, or superluminal probability, corresponding to negative evidence (retractions), and leading to split-complex probability and tachyonic or spacelike states.

bBernoulli pPhotons

cConsider bBernoulli observations as trials, with success as a unit of support and failure as a unit of opposition. aA success can be considered as 𝑑𝛼=1,𝑑𝛽=0$d\alpha=1, d\beta=0$ . tTherefore the change in epistemic spacetime interval is 𝑑𝑠2E=4𝑑𝛼𝑑𝛽=410=0$ds_{\mathrm{E}}^2 = 4 \, d\alpha \, d\beta = 4 \cdot 1 \cdot 0 = 0$. lLikewise, a failure would be 𝑑𝛼=0,𝑑𝛽=1$d\alpha=0, d\beta=1$ and also 𝑑𝑠2E=401=0$ds_{\mathrm{E}}^2 = 4 \cdot 0 \cdot 1 = 0$. iIn special relativity, an object with 𝑑𝑠2=0$ds^2 = 0$ is traveling at the speed of light, aka a null or lightlike trajectory, and only photons - massless particles - travel like this. tThus a trial is analogous to a photon.

nNext, let's take a look at the formula for epistemic mass 2𝛼𝛽$2\sqrt{ \alpha \beta }$ . tThis is the geometric mean scaled by 2, and might be interpreted as a measure of conflict. wWhen two photons collide, they can produce massive particles, and when opposing photons are treated collectively, they have an invariant mass. lLikewise, a belief gains epistemic mass when epistemic photons collide, which makes sense - a belief that doesn't lie at absolute certainty has some internal conflict. tTo generate mass, each success needs failures to "interact with" and vice versa.

iInterpreting dDistributions as bBeliefs

bBeta distributions describe fuzzy beliefs well when the beliefs are about propositions, due to the bBernoulli-bBinomial-bBeta connection and the tTrue-fFalse nature of propositions. tThe mode is the agent's single most likely degree of confidence (not the variance)! tThe variance is malleability or "sway-ability" of the belief. tThe mean is the "average belief on repeated draws", as if you elicited belief from an agent multiple times in either nearly the same conditions or multiple worlds. tThis merges the bBayesian and frequentist interpretations of probability. 50% is unsure, 99% is very sure the proposition is true, 1% is very sure it's false. tThe trials are how resistant an agent, component, or neuron is to changing its mind. mMany trials at 99% mode is "iI'm sure that iI'm sure that it's true" while few trials at 99% mode are "iI think it's true but iI would be easily swayed." mMany trials at 50% are "iI'm conflicted" or "iI'm sure that iI'm not sure" and few trials at 50% are nearly vacuous - "iI'm not sure at all."

dDirichlets would represent beliefs on multiple-choice questions. iI've also re-derived an obscure derivation to allow for nonparametric beliefs on number lines, making dDirichlet processes more flexible, though that's beyond the scope of this article. aAlso, a common misconception is that studies themselves have beliefs or provide evidence as beliefs - in my framework only agents hold beliefs, and beliefs about studies should be about a proposition about the study, such as a belief on "sStudy aA showed xX".

lLorentz fFactor

nNow that we have an evidence cone, we can relate probability to the lLorentz factor, typically defined as

𝛾=11(𝑣/𝑐)2
$$ \gamma=\frac{1}{\sqrt{ 1-(v/c)^2 }} $$

wWe can define an epistemic lLorentz factor

𝛾E=11𝑣2E
$$ \gamma_{\mathrm{E}}=\frac{1}{\sqrt{ 1-v_{\mathrm{E}}^2 }} $$

tTo relate back to the physical lLorentz factor, simply multiply the jJeffreys prior by 𝜋$\pi$:

𝐸𝑚𝑐2=𝛾=𝜋𝜛E
$$ \frac{E}{mc^2} = \gamma = \pi\varpi_{\mathrm{E}} $$

tThe epistemic lLorentz factor is related to several other equations. fFor example, fFisher information can be defined as

𝐼(𝜇)=1𝜇(1𝜇)
$$ I(\mu) = \frac{1}{\mu(1-\mu)} $$

cConvert 𝛾E$\gamma_{\mathrm{E}}$ back to normal probability coordinates

𝛾E=11(2𝜇1)2=114𝜇2+4𝜇1=14𝜇(𝜇+1)=12𝜇(1𝜇)
$$ \begin{aligned} \gamma_{\mathrm{E}} &= \frac{1}{\sqrt{ 1-(2\mu-1)^2 }}\\ &=\frac{1}{\sqrt{ 1-4\mu^2+4\mu-1 }}\\ &=\frac{1}{\sqrt{ 4\mu(-\mu+1)}}\\ &=\frac{1}{2\sqrt{ \mu(1-\mu)}}\\ \end{aligned} $$

tThen,

𝛾2𝐸=14𝜇(1𝜇)=14𝐼(𝜇)
$$ \begin{aligned} \gamma_{E}^2&=\frac{1}{4\mu(1-\mu)}\\ &=\frac{1}{4}I(\mu) \end{aligned} $$

aAs another example, consider a spring with normalized position between -1 and 1:

𝑥=cos𝜃
$$ x=\cos\theta $$

sSuch a spring spends its time as

𝑓(𝑥)=1𝜋1𝑥2
$$ f(x) = \frac{1}{\pi \sqrt{ 1-x^2 }} $$

Jeffreys, springs, E/mc^2

circle: uniform phase θspring: projected motion-10+1-10+1
n = 0
001π2time densityγE−10+1position x/A — also velocity vE
photos of the mass Jeffreys prior ϖE Lorentz factor γE/π — the same curve
$$x/A=\cos\theta,\qquad \gamma_{\mathrm E}=\frac{E}{mc^2}=\pi\,\varpi_{\mathrm E}(v_{\mathrm E})$$

wWhy 𝜋$\pi$?

iIf we use the lLorentz factor as a recentered probability density, we have to divide by the integral from 𝑣=𝑐=1$v=-c=-1$ to 𝑣=+𝑐=1$v=+c=1$.

1111𝑣2𝑐2𝑑𝑣=𝜋
$$ \int_{-1}^{1} \frac{1}{\sqrt{ 1- \frac{v^2}{c^2} }} dv = \pi $$

fFuture cContent

iIf you're inspired and understand the basics, maybe you could make content about your explorations in epistemic physics!

mMultidimensional eExtension by aAnalogy

sStart with a 1dD bBeta distribution. tThe null directions are 1$-1$ and 1$1$, or 𝑌$Y$ and 𝑁$N$, since in 𝑆0$S^0$ there are only two possible spatial directions. tTo extend the spatial bBeta naturally, start by assigning directions to the categories of a dDirichlet distribution. fFor example, for 𝐾=4$K=4$, the four categories could be nNorth, sSouth, eEast, and wWest.

tThere is a problem however - this dDirichlet distribution forms a diamond where velocities lie, and it excludes 36.3%$36.3\%$ of valid velocities, which can lie in a circle. iIncreasing 𝐾$K$ leads to more coverage, but to achieve 100%$100\%$ coverage, we need to take 𝐾$K$ to infinity, thus leading to a directional dDirichlet process.

aA similar construction to the jJeffreys prior still works here - a prior concentration of 1 on the dDirichlet - both are 1/𝐾$1/K$ per category, and imagine cutting the disk in half; each half carries 1/2$1/2$, matching Beta(1/2,1/2)$\operatorname{Beta}(1/2,1/2)$. gGoing to 3-dD and considering spatial, time, and complex dimensions is beyond the scope of this article.

oOther pPotential tTopics

  • mMore connections and predictions
    • sSpecific topics - klKL divergence, bBhattacharyya coefficient, hHellinger distance, amAM/gmGM inequality, the gGudermannian, aicAIC, time vs spatial dimensions, photon rocket science, mMoran process, and way more.
    • wWide variety of fields, especially sub-fields in physics, statistics, information theory, and philosophy, but there are even connections with fields you might not expect, like music.
  • iIndependence of evidence, fuzzy logic, opinion pooling and probability fusion, and bBeta mixtures
  • gGamma-distributed evidence, and the lLibby-nNovick, gGauss-hypergeometric, and gGeneralized bBeta pPrime distributions
  • gGroup theory
  • wWave-like and quantum beliefs
  • eExpansive speculation for fun
  • dDeriving a cCauchy "mean and variance" and showing equivalence with mMccCullagh's parameterization. yYou can take a look at the aiAI slop version here (iI curated the demo though).

fFun pPredictions

hHere's a small selection of potential predictions; iI hope to expand this list + details in future content. aA lot of these can be further formalized with equations.

  • iIn natural selection, curvature of rapidity should correspond to frequency-dependent selection or changing environmental conditions.
  • cConfirmation bias - as confidence increases, the diversity of sources an agent samples reduces as 1/𝛾=1𝑣2E$1/\gamma=\sqrt{ 1-v_{\mathrm{E}}^2 }$
    • fFor multidimensional beliefs, photons against a belief are blueshifted and compressed into an angle 1/𝛾$1/\gamma$, while photons confirming a belief are redshifted and spread out. tTherefore highly confident agents view opponents as a small, loud, homogeneous group while viewing supporters as diverse with many positions.
  • eEcho chambers, or polarized communities, have a cCurie temperature - adding independent evidence shifts the cCurie point while shared (dependent) evidence doesn't, though the direction of shift depends on evidence balance.
    • hHysteresis appears - the entrance to and exit from polarized states are different.
    • lLet older evidence be forgotten at some rate. iIf we model social amplification, independent evidence, and forgetting rate, we can predict how much independent evidence is needed to eliminate polarization, as well as the echo chamber recovery time.
  • mMotivated reasoning - say an agent weighs incoming evidence by its current belief: 𝑒𝜅𝑣$e^{\kappa v}$ for a success, and 𝑒𝜅𝑣$e^{-\kappa v}$ for a failure. mMotivated reasoning is harmless, in the long run, as long as 𝜅<1$\kappa<1$, but if 𝜅>1$\kappa>1$, the agent becomes overly confident.
  • oOrder effects in 2dD or higher - consider two beliefs aA and bB. aA-then-bB vs bB-then-aA should have wWigner rotation 𝑅12𝜙𝐴𝜙𝐵sin𝜃$R\approx\tfrac12\phi_A\phi_B\sin\theta$.

sSpeculation

  • tThere might be a rRindler horizon and uUnruh effect affecting beliefs - perhaps there's a boundary where evidence can no longer catch up to a changing belief, and rapidly changing beliefs see illusory weak evidence as uUnruh radiation.
    • rapidly persuaded agents may see more patterns in ambiguous evidence as compared to gradually persuaded agents.
    • an epistemic black hole may be created if current belief influences what kinds of evidence is possible to receive, or with other models of information dynamics.
  • tThe figures in this paper about wWeyl fermions, especially the correspondence between spacetime and energy-momentum cones, seems related. nNot that iI really know much about wWeyl fermions.

pPotential sSources and cCitations

tThis is not an academic paper, but iI wouldn't mind developing this section further. fFeel free to suggest corrections to my article, as well as potentially useful citations and sources (iI'll credit you)!

a little conversation

Leave a thought, ask a question, or select words in a paragraph to reply to that passage.

No published comments yet. You’re welcome to start the conversation.

Sign in to join the conversation