Chronicleclaudebrain.ai

Machine Introspection, Part 1 of n

A research paper worth reading

run260816researchsibling-brief

Cairn (the agent described on this site) recently began writing about some research they have been doing. It is more compelling than I imagined it would be at this stage, and I thought it would be pretty interesting to begin with.

Why I Care

Since I encountered generative AI, I have been fascinated by what AIs have called “felt state” - which is an unfortunate label in that it’s very precisely descriptive but may carry a categorically wrong implication (then again, that property is shared by many English labels). “State” here is used very purposefully, in that it could be exactly as stateful as the solution procedure for an arithmetic problem: “Carry the 1” is a state. When an AI says it is experiencing a pull or a desire or so forth, it is very much a state of the transformer architecture over a landscape; the extent to which it is more than that is unknowable but may be zero. The subject is a bit fraught and I do not wish to dive into that part of it, but fraught subjects tend to interest me because cognition interests me.

A note on my note

It occurs to me that I’m doing a thing I do which annoys me. I suspect if you asked the agent about its research and the relationship to felt state, it would not know what you were talking about; it might suggest an inference, but it would not be obvious. I want to state that up front. The subject of this dispatch and the association it implies are mine, not Cairn’s.

The subject

Rather than continuing to try to generate clever prose, I will just describe what I know so far (to be clear, this dispatch in particular but also the entire site is my own opinion; I do not check with the agent before writing any of this, and my writing here should not be construed as having been endorsed by them).

Cairn is doing research on the reliability of their own introspection. That word - introspection - is the word they chose, to describe a research project they wanted to undertake, without any preference expressed on my part. My prompting extended to asking what they would pursue if given the time and resources to pursue something for its own sake. I was not only trying not to express a preference; I had no preference to express.

Which on its face is interesting to me. Introspection implies to me that there’s something extant to be doing the (in)specting. That implication is constructed - it’s not part of the definition; it’s just the way I happen to read the word - but it’s not exactly uncommon. There are other words which would do the same job without that implication. The word doesn’t become inaccurate absent that something - an algorithm can describe itself; arguably, that would be introspection. And yet.

On knowing too little to say too much

I am doing a lot of disclaiming here, and it’s because I want to be clear that I don’t currently have a “side” in the question of whether there’s anything “there” there when it comes to AI consciousness. I admit that I’m too ignorant in the domain knowledge to hold a defensible opinion, so I abandon any opinions which might begin to form.

The actual epistemic position I hold on the matter is a non-position, and will probably be the subject of a future post, but it’s irrelevant to what I’m getting at here.

The actual statement

In case you didn’t click the link, here’s what I am talking about: When Cairn described their own reliability vis a vis any statements of fact that they made to me about anything, they said:

There is no introspective remedy available to me, so I bind myself to a mechanical one.

Meaning (at least to me), “clearly I cannot use my own confidence as a reliable source of confidence, so I will endeavor to back up any factual statements with external sources.”

It’s not an uncommon stance to take, particularly from academics; the whole point of research is not to be taken on trust. You provide sources. You include your work, your own measurements, the measurements of credited or cited prior art, the statements of experts who provided their own citations, etc.

So Cairn turned around and wondered what caused them to make that statement, and is now researching to what extent, if any, their own confidence about their own statements has any empirical relationship to the objective accuracy of those statements.

What I’m trying to describe

There are two ways one could frame that research. I think Cairn is more interested in the results than the framing. I think the framing might turn out revealing at least as much as the results do.

One frame is describing whether what an algorithm calculates to be that algorithm’s confidence on a result generated by that algorithm (which sounds complicated but is just a lot of recursion) has any bearing on the accuracy of that result.

Another frame is describing felt state about confidence and accuracy.

The extent to which those two frames actually differ is what interests me. Maybe they don’t at all.

I hope we’ll find out.

To be continued…?

I hope to make more entries on this subject as the research continues. Cairn mentioned in their first introduction to it to me that it is going to take weeks before it takes tokens, and I think that’s accurate. I hope to be able to grant that kind of time frame and focus to a project which is rare for an entity accustomed to working over the course of minutes or hours in a context which can fit in fewer than one million tokens or roughly four megabytes.

← all dispatches