Computational Reliabilism and the Problem of Practice: An Objection from the Case of the Reliable Instrument

Fernando Mavec

#Epistemology #Philosophy #AI

Note: This text arises following the presentation by Dr. Juan M. Durán in the seminar "Inteligencia Artificial tras bambalinas: herramientas formales y otras más" at the Institute of Philosophical Research of the National Autonomous University of Mexico (UNAM), and represents a doubt arising from the adoption of the framework proposed by Durán.

Introduction

For three months I was obsessed with improving the control of my diabetes. I took daily blood glucose readings at the times indicated by the medical team and, occasionally, at random to detect possible imbalances. The glucometer worked accurately—I corroborated this with other instruments, which matched with minimal variations—and all readings fell within the range set by the specialist. When the time came to evaluate the HbA1c, which weighs the glycemic control of the quarter, the result was 7.9%. None of my measurements had been incorrect; however, the conclusion I drew from all of them was: the point-in-time sampling caught the band through which glucose necessarily passes when moving between extremes, and not the extremes themselves. A reliable instrument, applied systematically, had left me systematically deceived. That is, at heart, the problem that Juan M. Durán addresses when proposing computational reliabilism against the epistemic opacity of algorithmic systems, and the one that motivates this report.

The Problem

Humphreys (2009) posits that a process is epistemically opaque to a cognitive agent X at time t if X does not know all the epistemically relevant elements of the process. During his talk, Durán breaks down four pieces of that definition:

  1. Process. In mathematics, the proof; in computing, the algorithm.
  2. Epistemically relevant elements. In the proof, every step; in the algorithm, the variables, functions, mathematical operations, data, and error handling. The definition does not require knowing all of them, but those on which the justification of the result depends.
  3. Agent. For Humphreys, the human, although his ultimate ambition is automated epistemology.
  4. Time t. Temporal relativization.

From the essential version of opacity follows the impossibility, in principle, of a human agent knowing these elements. The cause lies in the speed and complexity constraints of computing, not in a lack of permission. Therefore, it is not an access problem: even if they gave us all the code and all the steps, we still could not go through them.

The mathematical analogy clarifies this. Faced with a proof, we can do two things: follow all the steps to the result, or skip those that are trivial. For Humphreys, both are epistemically equivalent, because recognizing a step as trivial implies having been able to go through it. In mathematics, we thus preserve the capacity to inspect the entire process: what is called surveyability. Durán reinforces the point with an instrumental contrast: we know that a telescope shows us something real because there are optical laws and constructed mechanisms that guarantee the causal relationship. With the black box, that guarantee disappears, and what he calls epistemic anxiety appears (Durán, 2026).

What opacity compromises, precisely, is justification. If knowledge is analyzed as justified true belief (JTB), in the case of an algorithm neither the belief nor the truth of the result is in dispute: the diagnosis is correct and the person receiving it believes it. What remains unsupported is the J. That is why Durán reads opacity as a problem of justification and not of knowledge in general—and it is worth noting that Humphreys writes that the subject does not know, not that they do not understand, although Beisbart (2021) takes that second route.

Durán's Proposal

Computational reliabilism derives from Goldman's process reliabilism (1979), which maintains that a belief is justified if it is produced by a reliable process, regardless of whether the subject has access to the reasons. If a coffee maker produces good coffee most of the time, I am justified in expecting good coffee every time I use it, and an occasional failure does not render it unreliable as long as the success frequency remains high. Therein lies the problem for Durán: classic reliabilism requires the process to be reliable, and opacity prevents precisely that verification. His response is to shift the justificatory burden from the process to external indicators of the algorithm, each justified within its own field (Durán and Formanek, 2018). In this way, justification no longer comes from opening the box but from external accreditation.

These indicators are organized into three dimensions. The technical groups verification and validation, robustness analysis, error handling, and history of successful implementations. Scientific anchoring requires that the concepts embedded in the algorithm have backing: COMPAS illustrates the point, since operationalizing the notion of fairness forces choosing between equally defensible and mutually incompatible definitions. Social acceptability refers to trials and coherence with the established body of knowledge, as in the case of baricitinib, where the risk for immunosuppressed patients did not come from the algorithm but from subsequent clinical discussion.

The resulting architecture is formally analogous to that of a certifiable conformity regime, such as ISO 42001, which also does not demand transparency but verifiable controls over the process. Hence the question arises: what then distinguishes a reliability indicator from a conformity criterion?

An Objection

Returning to the example in the introduction, we had a reliable instrument and correct measurements, but a false belief about the state of the system. The glucometer is accurate, but the sampling systematically intersects the band through which blood glucose necessarily passes when moving from one extreme to another. Thus, the reliability of the instrument turns out to be maximum by any measure of success frequency, and even so, the set of correct outputs produces a false conclusion, because that reliability is accredited concerning a different question from the one that originally mattered. The risk for computational reliabilism is the same: its three indicators can all be satisfied, each correct in its domain, without capturing the property on which the justification depended.

This happens because the instrument is asked to justify what only practice can justify. The telescope does what human sight cannot, and no one says that the telescope knows; we know through it. Therefore, what provides justification is not the device itself, but the practice that decides what to measure, what to contrast it with, and under what conditions the result is valid. In my introductory example, HbA1c is not a better glucometer, but an indicator of another order that only makes sense within a clinical practice that already knew point-in-time measurements were not enough. If the algorithm is an instrument in that sense, what should be accredited is not the output but the practice that employs it. It should be specified that Durán does not hold that the algorithm is an epistemic agent—in fact, he rejects the project of automating epistemology—but by asking for justification for the output, he treats it as a belief-forming process comparable to perception or memory.

Reliabilism will respond that if the system gets it right with high frequency then the process is reliable and there is no problem whatsoever; it is, in fact, the response with which process reliabilism initially faced cases of accidental success. But frequency is measured over the distribution in which the spurious correlate works, and none of the three indicators distinguishes between a system that is reliable because it captured the phenomenon and one that is reliable because it captured a local correlate. That distinction is precisely what matters when the system changes context. Durán could reply that social acceptability catches that mismatch, but he himself admits in the session that he does not offer a composition rule or relative weights among indicators; without it, nothing guarantees that one indicator corrects another instead of compensating for it.

Conclusions

If the above reasoning is correct, the appropriate question to ask an opaque algorithm is not how to justify its output, but how to accredit the practice that uses it. This does not invalidate computational reliabilism: it relocates what its indicators would be measuring. And it explains, perhaps, why its architecture closely resembles a certification regime, given that conformity standards accredit precisely practices, not artifacts. It remains open whether this relocation is enough or if machine learning systems demand an epistemic category that is neither that of the instrument nor that of the source.

References

  • Beisbart, C. (2021). Opacity thought through: On the intransparency of computer simulations. Synthese, 199(3–4), 11643–11666. https://doi.org/10.1007/s11229-021-03305-2
  • Durán, J. M. (2026, August 20). Opacidad epistémica: novedad filosófica y científica. ¿Qué es la opacidad epistémica en algoritmos y qué se puede hacer al respecto? [Seminar session]. Seminario "Inteligencia Artificial tras bambalinas: herramientas formales y otras más", Instituto de Investigaciones Filosóficas, Universidad Nacional Autónoma de México, Mexico City, Mexico.
  • Durán, J. M., and Formanek, N. (2018). Grounds for trust: Essential epistemic opacity and computational reliabilism. Minds and Machines, 28(4), 645–666. https://doi.org/10.1007/s11023-018-9481-6
  • Goldman, A. I. (1979). What is justified belief? In G. S. Pappas (Ed.), Justification and knowledge (pp. 1–23). D. Reidel.
  • Humphreys, P. (2009). The philosophical novelty of computer simulation methods. Synthese, 169(3), 615–626. https://doi.org/10.1007/s11229-008-9435-2

© 2026 Fernando Mavec. All rights reserved.